Toronto Transit Commission · Subway · 1 Jan 2014 – 30 Jun 2026
It's not (just)
the trains
The subway's reliability problem is usually blamed on ageing equipment. Twelve and a half years of the TTC's own incident data say otherwise. Every tile above is a thousand minutes of service lost between 2014 and 2026. There are 606 of them. People cause 267. The trains only cause 50.
- Incidents that cost service time
- 86,269
- Minutes lost
- 605,922
- Caused by the trains
- 8.3%
Ride the line
Take the wall apart
Ask what breaks most often and the trains lead every other piece of equipment on the railway: 9,385 incident reports, more than track and tunnels, more than signals. That is the familiar story: ageing trains, tired equipment, a large maintenance bill.
Ask what costs the most service time and the trains come fifth. Below is every minute lost between 2014 and 2026, sorted by what was responsible. Incidents involving people passengers taken ill, disorderly behaviour, someone down at track level; cost the most amount of time.
That leaves a question. If the trains break down so often, why do they cost so little time?
Most incidents don't result in delays
If an incident is written up and no train is held, it costs no service time. Two thirds of everything the TTC logs is like that.
Trains are the clearest case: nearly three quarters of train incidents don't result in a delay, against a quarter of weather and outside-interference incidents. That is why the trains file more reports than track and tunnels do and still cost 34% less time than they do.
Each bar is one group's own hundred incidents, split by whether they cost any service time. The vertical line marks the rate across all incidents.
This does not mean the trains are in good condition. It means their failures are usually cheap. A maintenance budget aimed at rolling stock is aimed at eight percent of the lost time. Most of the rest is caused by people, which an engineering department cannot control.
That is the whole network. The next question is what it looks like on a single line.
The same record, in engineering units
The rail industry does not measure reliability in minutes lost. It uses two other numbers, and both make the same record look different.
The first is availability: the share of a line's running hours with nothing holding up service on it. Incidents that cost no time do not count against it. Line 1 scores 94%, which reads like a pass mark.
The missing six percent is 5,331 hours in which something on Line 1 was holding up service. That figure is measured, not assumed. Each incident is logged with a start time and the minutes it cost, so the two together mark out the window it covers, and windows that overlap are counted once. Availability hides its own size until it is converted back into time.
Availability, redrawn as time
4,564 days at 20 running hours. Worked out one line at a time.
Time between failures
The second number is how long a line runs before its next failure of a given kind. Pick a line below and the subsystems sort themselves by it.
The right-hand column, headed disruption, measures something different: how long service stayed disrupted each time. A subsystem can fail once a year and cost an hour when it does. Another can fail weekly and be cleared before you notice. Both numbers matter, and neither one alone shows where to spend.
The missing number: repair time
A reliability report normally carries a third number: how long the repair itself took, not how long passengers waited. That number is not in this data and is not published anywhere.
The disruption column is the closest available, and it measures something else: how long service stayed disrupted. A train with a jammed door is pushed out of the way and running again in minutes. The repair happens overnight in a yard.
The giveaway is that the figure barely moves. Across all 21 combinations of line and subsystem it sits between 5 and 15 minutes. Real repair times would vary far more than that. Clearing a passenger alarm and fixing a signal fault are not both ten-minute jobs. A band that narrow means the column is measuring recovery, not repair.
So far these calculations are for whole lines, averaged across twelve and a half years. But on which stations do these actually take place on?
Twelve and a half years, in four minutes
Here is a visual playback of what happened: every one of the
86,269
incidents that held back the trains. The bigger the dot, the more minutes it cost.
Try toggling off a subsystem to stop drawing it.
Press play, or drag the timeline. Click any station for what goes wrong there, its worst hour of the day, and how quickly it recovers. The Recovery speed tab recolours the map by the time it takes for a station to recover from delays.
Click any station
Busy stations recover fastest
The replay shows incidents piling up where the passengers are. Bloor–Yonge is almost constant. The quiet outer stations barely register.
That suggests the big interchanges are the weakest points on the subway. The Recovery speed tab shows they are not.
Average delay per station does not measure how well a station performs. It mostly measures what happens to go wrong there. A station near a hospital logs more medical emergencies, and those take a long time, so the station looks bad without being bad.
So each station is measured against itself. Take the causes logged there, look up what each of those causes costs across the network, and calculate their average. That is the delay the station should have had. Divide the delay it actually had by the average. A station that matches the network scores 1.00. Below 1.00 it clears incidents faster than the network does for the same causes; above 1.00, slower.
Measured that way, the ranking reverses. Sort all 70 stations by how many incidents they log: the quietest quarter sits at 1.11, the busiest quarter at 0.94.
It is a trend, not a rule. 4 of the busiest 17 stations still take longer than the network does for the same causes. Union, just behind them in size, is a fifth.
One thing the chart does not say: busy here means how many incidents a station logs, not how many passengers pass through it. The TTC does not publish per-station ridership over this period, so incident count stands in for traffic.
Hover or click any dot for its numbers
Woodbine is the clearest case. Nothing unusual happens there: disorderly passengers, door faults, people taken ill. It still takes about a third longer to clear them than the network takes for the same mix of causes. Bloor–Yonge has far more incidents and clears each one faster.
The likely reason is what a busy station has on hand when something goes wrong. Staff already there. Crossovers to run trains around the problem. A terminal nearby, so a train can be turned short. A quiet outer station has an incident and waits for help to arrive.
That changes the question. Reliability is not only about where equipment breaks. It is also about where the railway can recover quickly, and those are different places.
Everything so far treats twelve and a half years as a single block. That cannot answer the question people actually ask: is the subway getting better or worse?
Something changed in 2021
In 2021 the numbers get worse. The railway did not.
Take each part of the railway that has enough history to compare — every subsystem on every line, 20 of them — and compare the years before 2021 with the years after. 10 get worse, and 8 of those had been getting better until then. Trains, track, signals and outside causes all turn in the same year, on three lines with different equipment, different ages and different crews.
Equipment does not fail like that. Trains do not start breaking down more often in the same twelve months as tunnels, and Line 4 — five stations, its own trains — does not join in by chance.
What changed is how incidents were written down. Nobody outside the TTC can prove that from published data, but from 2021 the record says as much about TTC paperwork as it does about the railway.
One cause code shows how much damage a recording change can do.
PUOPO covers the platform door cameras an operator watches when
running a train without a guard. It cost service time about
15 times a year before
2021 and about 713 times
a year since, peaking at 855
in 2024. Nothing
suggests the cameras became
46 times worse.
That one code is most of Line 1's infrastructure trend. Leave it in and the failure rate climbs at 2.00 on the standard scale, where 1.00 is a rate that is not moving; take it out and the same track and tunnels sit at 1.10. Read without knowing this, the trend says the infrastructure is falling apart. It is not.
Treat any trend that crosses 2021 with suspicion, including the ones on this page.
That is a real limit on what any analysis of this data can claim, and it applies hardest to the strongest result the data has.
One real change was made to this railway during those years, and it came with something to compare it against.
Signalling fails on a schedule
Signalling causes 4.2% of the lost time. It gets a section anyway, because it is entirely the TTC's own equipment to maintain and replace, and because it keeps failing on a fixed schedule every year.
A solid ring is what an even year would look like. Signalling is worst in January and best in October, a 1.89× gap. The cause is the brutal winter: ice on the track circuits that detect where trains are, point machines freezing before they can move a switch, electronics stressed beside the track.
The TTC already works around this every winter, with switch heaters and de-icing trains. What the swing adds is a size for that work: it shows how much worse January is than the quietest month, and therefore how much should winter preparation be.
A pattern that arrives every January can at least be planned for. Predicting how long the next incident will last is a harder question, and there is a model for it.
What the model knows
A control centre faces a different question than any asked so far. An incident has just been reported, we know the cause, the station and the time. How long will service stay down? The answer decides whether to hold trains, turn them short, or wait.
A machine learning model was trained using the data to answer this question. Below is that model, running live. Set an incident with the three controls and it gives its estimate, with the range most incidents like it fall into.
Change the cause and watch the estimate move. Then change the station, and the time of day.
What it learned from, and what judged it
- 2014 to 2023Trained on this data.
- 2024 and 2025Tuned on this data.
- 2026, to 30 JuneJudged on this data
Predicted service disruption, in minutes
6.4minutes, typically
Eight in ten land between those ends.
Change a control to see how far it moves the estimate.
Set the cause to a speed-control fault and the model says about four minutes. Set it to a person struck by a train and it says about sixty-five. That one control covers the whole range; the station and the time of day only adjust what it lands on.
Take a disorderly passenger, the most common cause on the list. Across every station and hour the controls can reach, the estimate stays between 3.7 and 6.2 minutes: about 2.1 minutes of that is which station, and 0.9 of a minute is the time of day.
So the cause tells you what kind of incident you have. For a control centre deciding whether to hold trains, this is the most influential factor affecting their decision.
How much each input matters
Each input can be measured by scrambling it, so that it still looks like real data but carries no meaningful information, and then seeing how much worse the model's predictions get. A large drop means the model relied on that input. Almost no drop means it barely used it. Bigger numbers below mean heavier reliance.
So once you know what broke, knowing where and when it broke adds almost nothing. That was an open question for the TTC, and this is a clear answer to it.
The table nearly kept up
Every number below is an average miss, in minutes: give the thing an incident it has never seen, ask how long the delay will run, and this is how far out it typically is. Lower is better. The left column is the years used to tune the model; the right is the half-year held back from it.
| Average miss, in minutes | 2024–2025, tuning | 2026, held back |
|---|---|---|
| Machine learning model, tuned | 3.881 | 4.407 |
| Lookup table, by cause code | 4.013 | 4.534 |
The model was tested against the simplest possible rival: a lookup table. Find the cause code, take the delay that code usually runs, use that number. Two lines of arithmetic against a tuned machine learning model.
The tuning years are the ones used to adjust the model.
The 2026 data was held back to test the model.
The model is ahead on both: by
3.3% on the tuning years
and 2.8% on the held-back
one; about a tenth of a minute each time.
Three percent, for a tuned model against two lines of arithmetic. That is the finding, and it is barely a win. Anyone deciding whether to run a model here should weigh three percent against the cost of owning one. The table is the cheaper thing to own — anyone can read it, and leaving it un-refreshed for two years costs it four thousandths of a minute — but it is not free. It still needs rebuilding as the record grows, and when a cause code turns up that it has never seen, it quietly answers with the network median instead of admitting it does not know.
No better model fixes this, because the information is not in the data. Whether a supervisor reached the platform quickly, whether paramedics were already nearby, whether a spare operator was free; those decide how long an incident actually runs, and none of them are recorded anywhere.
The same shortage has now appeared in every section. It is time to name it.
End of the line
Two missing numbers
Twelve and a half years of the TTC's own data, and the most useful finding in it is not a number. It is that the investigation kept hitting the same wall from different directions.
The record says what was logged. It never says what happened next.
- Nobody can tell a slow repair from a slow recovery. Repair time is not published, so how long service stayed disrupted stands in for it. Those are different things, and the difference between them is how well maintenance is performing.
- The most useful finding here rests on a stand-in. Busy stations clear incidents faster, which means reliability depends partly on how quickly a station recovers, not only on how often it fails. Measuring that properly needs per-station ridership, which is not published either.
- Half the trends reverse in a single year, across unrelated equipment. Until the TTC says what changed in 2021, no trend on this page can be read as the railway physically wearing out.
- A tuned model beats two lines of arithmetic by about three percent. That ceiling is set by the data, not the method. No amount of modelling recovers information that was never written down.
Four findings, four methods, one cause. The TTC records the moment something goes wrong in detail, and almost nothing about what happens afterwards. Everything this page could not settle sits in that gap.
Publish restoration time. Explain what changed in 2021.
Neither is a large request. Neither needs a new system built, and neither is commercially sensitive. Both already sit in the TTC's own operational records. They are not in the file that gets published every month.
With them, anyone outside the TTC could tell a slow repair from a slow recovery, and say whether the railway is getting better or worse. That is the question everybody was asking in the first place. Without them, a page like this one is the best that public data allows, and it is not enough.