A win probability for every game, not just the ones we'd bet. Our number sits next to the market's own vig-free number so you can see exactly where we disagree — and every one of these gets graded, whether it became a play or not.
The bar is the whole idea. Where the two shaded halves meet is our win probability; the pale vertical tick is where the market put the same game. The distance between them is the disagreement, and on most nights it's barely visible — which is the honest picture. A dozen books pricing a baseball game land very close to where any decent model lands. A card flagged in amber is one we treat as our error rather than an opportunity: the market is not wrong by twelve points, and a model that thinks it is has a bug. Those are never eligible to become plays.
When we say a team wins 60% of the time, does it? This is the only claim on this site that can be checked in weeks instead of years — because every rated game counts, not just the handful we bet. Fifteen games a night is 450 data points a month; a two-play card takes over a decade to say anything at all.
No rated game has finished yet. Once they start grading, this line reports how far off our probabilities were.
An Elo rating built by walking every completed game in date order. It needs no training and no fitted coefficients — it calibrates itself against what actually happened, and a rating never contains information from a game that hadn't been played yet.
A season-to-date ERA gap between the two starters, converted to win probability and capped so one disastrous outing can't swing a rating. This conversion is an estimate rather than a fitted number, and it's the most likely thing on this page to be wrong. The calibration table above is what will catch it.
Each card carries both teams' records and how that starter has historically fared against tonight's club, because they're the first things anyone wants to see. Neither touches the number. A record is already inside the rating — counting it twice is double-counting — and a career line against one club is usually a dozen innings of luck, which is why any sample under three starts is greyed out and labelled thin.
A team's last ten games are already inside its rating, and weighting them again double-counts the same information while adding noise. Hot and cold streaks are mostly variance wearing a narrative.
The expected result is that the market beats us. Closing prices are set by people with better models, better data and money at risk. Publishing a number that gets scored against theirs isn't a claim to be sharper — it's a claim precise enough to be proven wrong, which is rarer in this industry than being right.
A rating on every game doesn't mean a bet on every game. Most nights the market and our number agree closely enough that there's nothing to do.