How we measure forecast quality, and why we publish it in the open.
When the model says 60%, does it happen 60% of the time? Closer bars mean better-calibrated probabilities.
On 13 host games played in a host country (USA/Canada/Mexico), the model scored with the home edge applied vs a neutral-venue counterpart — same games, same model. Lower Brier = better.
So far the host edge has improved accuracy on these games (Brier 0.2224 → 0.2062).
Accuracy over the 98 games predicted by the current model version, separated from earlier games that used a prior version. The full record above is never rewritten — this just shows how the current model is doing on its own forecasts.
On 101 games with a usable pre-kickoff market line, the model scored against the sharp benchmark on identical fixtures — a like-for-like skill measure that controls for tournament variance (an upset-heavy tournament punishes the market too). Lower Brier = better.
So far the model is behind the sharp market on these games (Brier 0.2541 vs 0.2330) — the gap is published either way.
Brier Score is the standard measure of probabilistic forecast accuracy. It is the mean squared difference between the probabilities a model assigned and what actually happened. Lower is better; a perfectly calibrated, confident model scores near zero.
Crucially, it rewards calibration, not bravado. A model that confidently calls the wrong outcome is punished more than one that expressed appropriate uncertainty.
Most prediction products never tell you how accurate they are. We take the opposite position: every forecast is scored, and the running Brier Score is published per competition and updated continuously during tournaments.
This is the difference between accuracy that is asserted and accuracy that is accountable. If our models drift, the score shows it before we do.