PARALLAXEDGE
How It Works / Accuracy & Brier Score
ModelsAccuracyExplainabilityData

Accuracy & Brier Score

How we measure forecast quality, and why we publish it in the open.

Live · WC2026
0.2531
Running Brier score over 103 scored matches
▼ 0.0569 vs baseline
65%
Outcome hit rate
0.3100
Backtest baseline
103
Matches scored

Calibration

predicted vs. observed

When the model says 60%, does it happen 60% of the time? Closer bars mean better-calibrated probabilities.

13%
11% obs.n=70
28%
22% obs.n=152
49%
66% obs.n=50
69%
78% obs.n=32
84%
60% obs.n=5
predicted observed

Host-nation advantage

with vs without the home edge

On 13 host games played in a host country (USA/Canada/Mexico), the model scored with the home edge applied vs a neutral-venue counterpart — same games, same model. Lower Brier = better.

Brier
0.2062 w/ edge
0.2224 neutral
Outcome hit rate
85% w/ edge
77% neutral

So far the host edge has improved accuracy on these games (Brier 0.2224 → 0.2062).

Current model

since host-advantage fix

Accuracy over the 98 games predicted by the current model version, separated from earlier games that used a prior version. The full record above is never rewritten — this just shows how the current model is doing on its own forecasts.

Brier
0.2490 current
0.2531 full record (103)
Outcome hit rate
66% current

Model vs the market

Pinnacle · same games

On 101 games with a usable pre-kickoff market line, the model scored against the sharp benchmark on identical fixtures — a like-for-like skill measure that controls for tournament variance (an upset-heavy tournament punishes the market too). Lower Brier = better.

Brier
0.2541 model
0.2330 market
Outcome hit rate
65% model
68% market

So far the model is behind the sharp market on these games (Brier 0.2541 vs 0.2330) — the gap is published either way.

Updated Jul 19, 10:03 PM UTC

What Brier Score measures

Brier Score is the standard measure of probabilistic forecast accuracy. It is the mean squared difference between the probabilities a model assigned and what actually happened. Lower is better; a perfectly calibrated, confident model scores near zero.

Crucially, it rewards calibration, not bravado. A model that confidently calls the wrong outcome is punished more than one that expressed appropriate uncertainty.

Why we publish it

Most prediction products never tell you how accurate they are. We take the opposite position: every forecast is scored, and the running Brier Score is published per competition and updated continuously during tournaments.

This is the difference between accuracy that is asserted and accuracy that is accountable. If our models drift, the score shows it before we do.

Want deeper insights?
Join the waitlist for advanced simulations, full distributions, and model explanations.
Join the Waitlist