MethodologyMéthodologie
What the estimates are made of, and how wrong they are

Data

Every estimate is deterministic: no AI text, no live flight API bills. Three public data sources feed the model: the FAA's SWIM/TFMS feed (live flight plans, departures, and cancellations for U.S. airspace, consumed continuously), 3+ years of U.S. DOT Bureau of Transportation Statistics on-time records (the training ground truth), and National Weather Service forecasts for both ends of your route.

Model

A gradient-boosted tree model with monotone constraints: features with a known direction (storms, a disrupted origin airport, a carrier having a bad day) are only allowed to push risk up. That constraint costs a little accuracy and buys reasons that always point the right way. Probabilities are calibrated, and serving mirrors training's imputation exactly; a sanity gate blocks any deploy where a carrier scores outside sane bounds on a plain row.

Verified accuracy

Backtested on 25,000 real flights (August 2026) through the exact serving path, using only what was knowable before departure: flights stamped not cooked were cancelled 0.7% of the time, preheating 1.6%, sizzling 4.8%, cooking 12.7%, and cooked 40.5%. Every tier cancels more often than the tier below it. Ranking quality (ROC-AUC) was 0.79. Absolute probabilities ran somewhat hot against the FAA feed's cancel definition; recalibration waits on the next DOT ground-truth release.

The delay estimates are backtested too (8,600 flights): ranking quality 0.76, and the calibration runs honest to conservative. When the board says a flight has a 40%+ chance of running 30+ minutes late, such flights actually ran late more than half the time; flights called safe were late about 7% of the time.

Estimates are statistics, not travel advice; a cooking flight usually still flies.

← Check a flight · Baked Airports