Architecture Document — CTA Public Transport Optimisation System
Version: 1.0 Date: 2026-03-12 Status: Baselined Standard: 4+1 Architectural View Model (Kruchten, 1995) Notation: ArchiMate 3.1 concepts rendered as Mermaid diagrams
Table of Contents
- Document Purpose and Scope
- Architectural Drivers
- Use Case View (+1)
- Logical View
- Process View
- Development View
- Physical View
- Architectural Decisions Summary
- Risks and Technical Debt
1. Document Purpose and Scope
1.1 Purpose
This document describes the software architecture of the Chicago Transit Authority (CTA) Public Transport Optimisation System. It is structured according to the 4+1 Architectural View Model (Kruchten, IEEE Software 1995), which organises the architecture into five complementary views, each addressing the concerns of a different stakeholder group:
| View | Primary Audience | Central Concern |
|---|---|---|
| Use Case (+1) | All stakeholders | Scenarios that drive architectural decisions |
| Logical | Architects, developers | Functional decomposition and key abstractions |
| Process | Architects, integrators | Concurrency, data flows, runtime behaviour |
| Development | Developers, build engineers | Module structure, package organisation |
| Physical | Operations, DevOps | Deployment topology, infrastructure mapping |
Diagrams use Mermaid syntax and follow ArchiMate 3.1 layering conventions:
- Technology Layer — infrastructure elements (brokers, databases, containers)
- Application Layer — software components and their interfaces
- Business Layer — business processes and actors that the system serves
1.2 System Overview
The system is a real-time streaming pipeline that ingests simulated operational data from the CTA elevated rail network ("L"), processes it through multiple transformation stages, and presents a live transit status dashboard. It demonstrates a full Event-Driven Architecture (EDA) on the Confluent Kafka platform.
1.3 Scope
- Three train lines: Blue, Red, Green (each with 10 trains, bidirectional)
- Station arrival events, turnstile ridership counts, and weather telemetry
- Static station reference data from PostgreSQL
- A browser-accessible real-time status dashboard
2. Architectural Drivers
2.1 Quality Attribute Requirements
| ID | Quality Attribute | Scenario | Architectural Response |
|---|---|---|---|
| QA-01 | Throughput | 3 lines × stations × 10 trains produce arrival events every 5 s | 10-partition Kafka topic; AvroProducer batching |
| QA-02 | Decoupling | New consumers must not require producer changes | All communication via Kafka topics (no direct calls) |
| QA-03 | Schema Evolution | Fields may be added to events over time | Avro + Schema Registry with compatibility enforcement |
| QA-04 | Replayability | Dashboard must recover state on restart | Consumers start from offset_earliest; Faust rebuilds table from log |
| QA-05 | Responsiveness | Dashboard must serve HTTP requests without stalling Kafka polling | Tornado async IO loop; consumers as coroutines |
| QA-06 | Extensibility | Station reference data changes without code deployment | Kafka Connect JDBC connector; consumers subscribe to topic |
2.2 Constraints
- Python-only application code (no JVM services authored in-house)
- Single-host Docker Compose deployment (development / demonstration environment)
- Confluent Platform 5.2.2 (fixed version)
3. Use Case View (+1)
The Use Case View captures the key scenarios that motivated and validate the architectural decisions. In the 4+1 model this view acts as the glue — each scenario exercises a slice through every other view.
3.1 Actor Diagram
3.2 Key Scenarios
UC-01 — View Live Transit Status
Trigger: Transit operator opens http://localhost:8888
Flow: Tornado serves status.html populated from in-memory Lines and Weather state
that is continuously updated by four Kafka consumers running as async coroutines.
Architectural relevance: Drives the Tornado async server choice (ADR-006) and the
requirement for in-process Kafka consumer coroutines.
UC-02 — Publish Train Arrival Event
Trigger: Simulation time step advances; a train moves to the next station.
Flow: Station.run() → AvroProducer.produce() → Schema Registry validates Avro →
Kafka topic org.chicago.cta.station.arrivals.t001 → KafkaConsumer in server →
Lines.process_message() → UI state updated.
Architectural relevance: Establishes the end-to-end Kafka + Avro pipeline (ADR-001, ADR-002).
UC-04 — Publish Weather Reading
Trigger: Simulation hour boundary.
Flow: Weather.run() → HTTP POST to Kafka REST Proxy → Kafka topic
org.chicago.cta.weather.v1 → KafkaConsumer in server → Weather.process_message().
Architectural relevance: Demonstrates the REST Proxy integration path (ADR-005).