70% isn't a yes
A single probabilistic call can't be graded true or false on its own. A 70% forecast that misses can still have been a good call — and a lucky 70% hit doesn't prove skill.
A probability isn't a verdict on one event. It's a claim you can only judge across many.
A single probabilistic call can't be graded true or false on its own. A 70% forecast that misses can still have been a good call — and a lucky 70% hit doesn't prove skill.
Calibration asks a different question: across every time a forecaster said "70%," did roughly 70 in 100 of those things happen? One event proves nothing. A hundred start to.
Introduced in 1950 to grade U.S. Weather Bureau forecasters, it scores a whole set of probability calls from 0 (perfect) to 1 (worst). A blind coin-flip on 50/50 events averages 0.25.
A forecaster is well-calibrated only if its 70% calls land about 70% of the time and its 90% calls land about 90% — visible only in an open, gradable record kept over time.
Sources: Brier, G. W. (1950) — Verification of Forecasts Expressed in Terms of Probability, Monthly Weather Review · Metaculus — public forecaster track record and calibration data · NFNO Lab — publishes probability, resolution criteria, and grade for every forecast