GIS Analyst · Chicago

GIS that makes the overlooked undeniable.

Real maps, real data, real people who needed someone to notice. New ones posted here roughly every week.

← Back to civic work
Civic — Independent Study

The Buses That Never Came

A ghost bus is a trip that exists on the schedule and never appears on the street. Across sixteen US transit systems, 5.8% of scheduled bus trips left no trace in the agency's own realtime feed.

16 transit systems, 12 service days 1,359,631 scheduled trips checked 78,966 never observed
Ranked bar chart in coral on near-black, showing ghost bus rates for sixteen US transit systems from San Antonio at 0.4 percent to Chicago and Baltimore at 14.1 percent.
16Transit systems
12Service days
1.36MScheduled trips checked
78,966Never observed

What was measured

Every transit agency publishes two things: a schedule saying which trips are supposed to run, and a realtime feed reporting where its vehicles actually are. Compare them and you can count the trips that were promised but never showed up.

From 3 to 15 September 2026 I collected vehicle positions every 90 seconds from 19 US agencies, alongside a daily snapshot of each agency's published schedule. For each service day I took the set of trips scheduled to run and checked which of them ever appeared in the realtime feed. A scheduled trip that is never seen is counted as a ghost.

That is a deliberately generous test. A trip counts as having run if it reported its position even once, so a bus that appeared for two minutes and vanished still counts as delivered. The real shortfall riders experience is larger than what follows.

The count

share of scheduled bus trips never seen in the realtime feed 0% 5% 10% 15% San Antonio San Antonio (VIA Metropolitan Transit): 0.4% ghost, 11 clean days 0.4% Denver Denver (RTD): 1.7% ghost, 11 clean days 1.7% Milwaukee Milwaukee (MCTS): 1.9% ghost, 9 clean days 1.9% Orlando Orlando (LYNX): 2.3% ghost, 9 clean days 2.3% Phoenix Phoenix (Valley Metro): 2.4% ghost, 11 clean days 2.4% Los Angeles Los Angeles (LA Metro): 2.7% ghost, 9 clean days 2.7% Charlotte Charlotte (CATS): 2.8% ghost, 11 clean days 2.8% Boston Boston (MBTA): 3.8% ghost, 11 clean days 3.8% New York New York (MTA New York City Bus): 4.9% ghost, 6 clean days 4.9% Philadelphia Philadelphia (SEPTA): 5.3% ghost, 10 clean days 5.3% Honolulu Honolulu (TheBus): 6.2% ghost, 11 clean days 6.2% Minneapolis Minneapolis (Metro Transit): 7.1% ghost, 11 clean days 7.1% Austin Austin (CapMetro): 7.6% ghost, 11 clean days 7.6% Pittsburgh Pittsburgh (PRT): 10.2% ghost, 11 clean days 10.2% Chicago Chicago (CTA): 14.1% ghost, 9 clean days 14.1% Baltimore Baltimore (MTA Maryland): 14.1% ghost, 9 clean days 14.1%
Share of scheduled bus trips never observed, pooled across all clean service days. Hover any bar for the agency and day count. Atlanta, Miami and LA Rail were collected but are not measurable; see the exclusions below.

The spread is wide. San Antonio's VIA delivered essentially everything it promised, at 0.4%. Chicago and Baltimore sit at 14.1%, meaning roughly one scheduled trip in seven left no trace at all. The median system lands at 4.3%.

Ranking cities against each other is the obvious thing to do with this table and also the thing to be most careful about. These are not equally comparable systems: they differ in fleet size, in how much of the fleet reports, and in how their feeds behave. What the number measures cleanly is a system against itself, day to day.

CityAgencyClean daysGhost rateWeekdayWeekend
San AntonioVIA Metropolitan Transit110.4%0.4%0.5%
DenverRTD111.7%2.0%0.8%
MilwaukeeMCTS91.9%1.2%4.1%
OrlandoLYNX92.3%3.2%0.7%
PhoenixValley Metro112.4%2.7%1.4%
Los AngelesLA Metro92.7%2.6%2.9%
CharlotteCATS112.8%2.8%2.8%
BostonMBTA113.8%3.4%4.7%
New YorkMTA New York City Bus64.9%5.0%4.2%
PhiladelphiaSEPTA105.3%5.3%5.5%
HonoluluTheBus116.2%7.2%3.4%
MinneapolisMetro Transit117.1%8.6%3.1%
AustinCapMetro117.6%7.7%7.5%
PittsburghPRT1110.2%10.4%9.5%
ChicagoCTA914.1%13.7%14.7%
BaltimoreMTA Maryland914.1%10.8%20.8%
Clean days are service days that survived all four validity tests. Weekday and weekend columns are computed from the same pooled trips.

Weekday against weekend

Most systems run better on weekends, which is less surprising than it first sounds: there is far less service scheduled, so there is more slack to deliver it. Denver, Orlando, Minneapolis and Honolulu all improve noticeably once the weekday peak disappears.

Baltimore goes the other way, and does so hard.

Weekday Weekend
weekday and weekend ghost rate, same scale 0% 5% 10% 15% 20% San Antonio San Antonio weekday: 0.4% San Antonio weekend: 0.5% Denver Denver weekday: 2.0% Denver weekend: 0.8% Milwaukee Milwaukee weekday: 1.2% Milwaukee weekend: 4.1% Orlando Orlando weekday: 3.2% Orlando weekend: 0.7% Phoenix Phoenix weekday: 2.7% Phoenix weekend: 1.4% Los Angeles Los Angeles weekday: 2.6% Los Angeles weekend: 2.9% Charlotte Charlotte weekday: 2.8% Charlotte weekend: 2.8% Boston Boston weekday: 3.4% Boston weekend: 4.7% New York New York weekday: 5.0% New York weekend: 4.2% Philadelphia Philadelphia weekday: 5.3% Philadelphia weekend: 5.5% Honolulu Honolulu weekday: 7.2% Honolulu weekend: 3.4% 3.8 pts worse, weekday Minneapolis Minneapolis weekday: 8.6% Minneapolis weekend: 3.1% 5.5 pts worse, weekday Austin Austin weekday: 7.7% Austin weekend: 7.5% Pittsburgh Pittsburgh weekday: 10.4% Pittsburgh weekend: 9.5% Chicago Chicago weekday: 13.7% Chicago weekend: 14.7% Baltimore Baltimore weekday: 10.8% Baltimore weekend: 20.8% 10.0 pts worse, weekend
Gaps of 3 points or more are labelled. Weekend figures rest on three service days (two Saturdays and one Sunday), so treat them as a signal worth following rather than a settled result.
The caveat that matters most

MTA Maryland's weekend ghost rate is 20.8% against 10.8% on weekdays, the largest weekday to weekend swing in the data. It is also built on three weekend days. I would not publish that as a finding about Baltimore yet. I am publishing it as the thing I intend to look at next.

Four ways this goes wrong

Every number above survived four tests. Each one exists because it caught something that had already produced a confident, wrong answer, and each catches a failure the others cannot see.

TestThresholdWhat it catchesExcluded
Schedule joinjoin_ok ≥ 90%Whether the two sides use the same trip identifiers at all. CTA publishes a Bus Tracker id, not a GTFS trip_id, so Chicago read as 41% realized before it was resolved on scheduled start time instead.3 city-days
Local service daymidnight to midnight, agency timezoneThe collector names its files by UTC date, but a service day is local. In Honolulu that put 46% of every file in the wrong day. It hid on weekdays, because consecutive weekdays reuse trip ids, then collapsed at every Friday-to-Saturday boundary.a correction, not an exclusion
Fleet coverage≤ 21 scheduled trips per reporting vehicleWhether the feed publishes the whole operating fleet. MARTA reports about 276 buses against 6,918 scheduled trips, the same 187 vehicles every day. Its unobserved trips are buses that ran without reporting.12 city-days, all Atlanta
Day completenessnull trip_id < 75%, no gap over 90 minutesWhether the day was actually captured. CapMetro spent 9 September publishing positions with no trip attached: 8.8% of rows carried a trip id, and the day scored a perfect join against 154 trips.22 city-days
160 city-days passed all four. 68 were excluded.

The Atlanta case is the one worth dwelling on, because it was almost the headline. MARTA came in at 37.4% ghost, stable across every day, with a 97.3% schedule join. Nothing in the join diagnostics objected. What gave it away was arithmetic: 6,918 scheduled trips against 276 reporting buses is 25.1 trips per vehicle, where every other system in the set falls between 3.5 and 15.0. No bus turns 25 trips in a day. MARTA's fleet is not failing to run, it is failing to report, and the two are indistinguishable from the outside unless you check.

A measurement that only has one way to be wrong is a measurement you should not trust. The timezone bug is the clearest example of why. It was invisible for the entire first pass: weekday joins sat at 99% because consecutive weekdays reuse the same trip ids, so the misattributed rows matched anyway. It only surfaced at service-type boundaries, where Friday and Saturday trip ids overlap by exactly zero.

What this does not show

It is not a punctuality measure. A trip that ran 40 minutes late and a trip that ran exactly on time are both counted as delivered. Only complete absence registers.

Reporting failures look like service failures. Atlanta was caught. A milder version of the same problem, an agency where 10% of buses do not report, would inflate that system's rate and would not trip the threshold.

Twelve days is a fortnight in September. There is no seasonal context here, no winter, no school calendar, and one of the twelve days was Labor Day.

The weekend sample is thin. Two Saturdays and one Sunday. Every weekend figure in this piece carries that qualification.

Three systems were collected but could not be measured. Atlanta for fleet under-reporting, Miami and LA Rail because neither has a schedule feed configured to compare against.

Method and data

Vehicle positions were polled every 90 seconds from each agency's GTFS-Realtime feed, or its native API where no GTFS-RT endpoint exists. Schedules come from a daily GTFS snapshot per agency, deduplicated by content hash, which matters because MTA and SEPTA both rolled a service pick mid-collection and publish only the current one. A scheduled trip set is resolved from calendar.txt and calendar_dates.txt for the local service date, restricted to bus and trolleybus route types, and matched against distinct trip ids observed in the local service day window including owl trips past midnight.

Collection is still running. This analysis covers 3 to 15 September 2026 and will be extended as the weekend sample deepens.

Sources

Every figure above comes from data the agencies publish themselves. Realtime endpoints are listed below with query strings removed; the feeds that require an API key were accessed with credentials held locally and not reproduced here. Static schedules were retrieved through the Mobility Database catalog, so each agency's schedule is cited by its stable catalog id rather than by a URL that moves.

CityAgencyRealtime endpointProtocolMDB id
San AntonioVIA Metropolitan Transitgtfs.viainfo.net/vehicle/vehiclepositions.pbGTFS-RT2348
DenverRTDopen-data.rtd-denver.com/files/gtfs-rt/rtd/VehiclePosition.pbGTFS-RT178
MilwaukeeMCTSrealtime.ridemcts.com/gtfsrt/vehiclesGTFS-RT2127
OrlandoLYNXgtfsrt.golynx.com/gtfsrt/GTFS_VehiclePositions.pbGTFS-RT347
PhoenixValley Metromna.mecatran.com/utw/ws/gtfsfeed/vehicles/valleymetroGTFS-RT147
Los AngelesLA Metroapi.goswift.ly/real-time/lametro/gtfs-rt-vehicle-positionsGTFS-RT29
CharlotteCATSgtfsrealtime.ridetransit.org/GTFSRealTime/Vehicle/VehiclePositions.pbGTFS-RT2265
BostonMBTAcdn.mbta.com/realtime/VehiclePositions.pbGTFS-RT437
New YorkMTA New York City Busgtfsrt.prod.obanyc.com/vehiclePositionsGTFS-RT510, 512, 513, 514, 520, 528
PhiladelphiaSEPTAwww3.septa.org/gtfsrt/septa-pa-us/Vehicle/rtVehiclePosition.pbGTFS-RT502
HonoluluTheBusapi.thebus.org/arrivalsJSON/TheBus API2350
MinneapolisMetro Transitsvc.metrotransit.org/mtgtfs/vehiclepositions.pbGTFS-RT205
AustinCapMetrodata.texas.gov/download/eiei-9rpfGTFS-RT150
PittsburghPRTtruetime.portauthority.org/gtfsrt-bus/vehiclesGTFS-RT409
ChicagoCTActabustracker.com/bustime/api/v3/getvehiclesCTA Bus Tracker389
BaltimoreMTA Marylandapi.goswift.ly/real-time/mta-maryland/gtfs-rt-vehicle-positionsGTFS-RT466
The 16 measurable systems. New York is six separate MTA borough feeds merged into one city. Two systems do not publish GTFS-Realtime: Chicago uses the CTA Bus Tracker API, and Honolulu publishes stop arrivals rather than vehicle positions, so its observations are derived from polling 140 stops.

Collected but not counted

These three systems were polled on the same schedule as the rest and are excluded from every figure in this piece.

CityAgencyRealtime endpointMDB idWhy it is excluded
AtlantaMARTAgtfs-rt.itsmarta.com/TMGTFSRealTimeWebService/vehicle/vehiclepositions368Fleet under-reporting: 25.1 scheduled trips per reporting vehicle.
MiamiMiami-Dade Transitapi.goswift.ly/real-time/miami/gtfs-rt-vehicle-positionsnoneNo static schedule feed configured, so nothing to compare against.
Los Angeles railLA Metro railapi.goswift.ly/real-time/lametro-rail/gtfs-rt-vehicle-positionsnoneNo static schedule feed configured.
Their raw collection continues, so Miami and LA Rail become measurable as soon as a schedule feed is added.

Software

Collection and analysis were written in Python with DuckDB for the schedule joins. Schedules are parsed from the GTFS static specification; realtime feeds are parsed as GTFS-Realtime protocol buffers except where noted above.

The collectors, the schedule snapshotter and the analysis are all in the repo, along with the four validity tests and the per-day results this page is built from. Credentials are read from an untracked .env file; the 13 agencies that need no key run out of the box.

View the code (opens in a new tab)