Reliability table builder
A forecaster that says "70%" should be right about 70 times in 100. Paste your forecasts and what happened, and this page builds the table that checks it.
How to read it
Each row groups the forecasts that claimed a similar chance. Average claim is the mean probability stated in that band; observed is the share of those forecasts whose outcome happened. When the two are close in every band the forecaster is well calibrated. A band where observed sits well below the claim is overconfident; well above, underconfident.
The Brier score is the mean squared gap between each probability and its outcome (1 or 0). Lower is better; always guessing 50% scores 0.25 on balanced data. It rewards both calibration and sharpness, so a forecaster can be perfectly calibrated and still have a poor Brier score if it never commits.
Small samples swing. With 20 forecasts a band may hold two or three cases, so its observed share can only be 0%, 33%, 50% or 100%. Treat anything under about 100 forecasts per band as a sketch, not a verdict.
A worked example on real data
The market calibration study ran this exact table on both sides of 18,238 settled sports markets. Every band landed within about two points of its claim, with outcomes priced below 50% happening slightly more often than claimed and those above 50% slightly less. Its open data set, one row per outcome with the margin-free probability, is a good file to paste here if you want to see what a large, well-calibrated sample looks like. Margin-free model probabilities to test the same way are available from the Bet Better API with no key.