Monthly Elo snapshots. Pick any teams below (up to 8) — each keeps its color while selected. A ★ marks a World Series, placed on that season's last rating: the Series ends in late October or November and the model stops at the regular season, so there is no later point to hang it on. Hover for the date it was actually won. scroll to zoom · drag to pan · double-click to reset
Elo as distance from the 1500 league mean, ordered by rating. Blue above average, red below. A franchise carries its current name at every date, so its history reads as one team. A ★ marks whoever went on to win that season's World Series — it shows for every month of the season, not just the ones after it was won, and a season still being played carries none. click a row to add or remove that team from the trajectories chart · drag the slider to read the table at any month in the data
What each model said against what happened. On the diagonal means honest: when it says 70% they win 70%, and when it says a 9-run game they score 9. Probabilities are bucketed on fixed 10% edges — the question is literally "when it says 70%" — and runs by decile of the forecast, since a runs forecast crowds around the league mean. A point above the line means the model said too little, below means it said too much. The bars underneath show where the forecasts actually land — each spans its own bucket, and its height is games per unit, not raw count. That distinction matters on the runs panels: those buckets are deciles, so every one holds a tenth of the sample and plotting the count drew ten identical bars. The middle deciles are narrow because that is where nearly every forecast sits, which is what the heights now say. Read them alongside the dots: a point far off the line above a very low bar is a handful of games, not a broken model. Panels run up the ladder — M1 through M5, and within each model p(home), total, margin. M3 is absent because it predicts a player's rate, not a game.
What each model is built from. Every layer was graded against the model before it; the rejected ones are still in the code behind flags, because a list of only the wins is a sales sheet. Scores on the right are re-graded at startup on one shared sample of games.