Mental model
Forecast Calibration
The degree to which predicted probabilities match actual outcomes over time—the difference between what you expect to happen and what actually happens.
Discover
A meteorologist says there's a 70% chance of rain every day for a week. It rains on 5 of those 7 days. Is this forecaster well-calibrated?
Test your intuition
Discover what calibration really means and why it matters.
Understand
Understand
Forecast calibration measures whether your stated confidence matches reality over time. If you say something is 70% likely, it should happen about 70% of the time when you make that prediction. Think of it like a shower dial: if you set it to "warm," the water should actually feel warm, not freezing or scalding. When predictions are well-calibrated, you can trust the numbers to guide decisions—from carrying an umbrella to investing resources. Check this: Track how often your 80% confident predictions actually come true.
Full explanation
Full explanation
Calibration separates overconfident forecasters from underconfident ones. A well-calibrated predictor who says something is 30% likely will see that event occur approximately 30% of the time across many similar predictions. It's not about being right or wrong on any single forecast—it's about the long-run alignment between probabilities and outcomes.
This matters because uncalibrated forecasts mislead decision-makers. If a doctor consistently says "there's a 10% chance of complications" but complications happen 40% of the time, patients can't make informed choices. Similarly, if a manager estimates "90% chance of on-time delivery" but only hits 60%, stakeholders learn to ignore the numbers entirely.
The good news: calibration can be measured and improved. By recording your confidence level (10%, 50%, 90%) for each prediction and tracking outcomes, you develop a personal calibration curve. Overconfident people see their 70% predictions materializing only 40% of the time; underconfident people see them happen 90% of the time. Awareness is the first step toward alignment.
Research
Research
Calibration research reveals a striking pattern: expertise sometimes correlates with calibration challenges; research shows mixed results depending on domain. This calibration-overconfidence paradox appears across domains—intelligence analysts routinely overestimate the certainty of their judgments compared to their actual track records [1]. Training can help: techniques like outcome feedback and considering reasons why predictions might be wrong reduce overconfidence and improve the alignment between stated probabilities and observed frequencies [1].
Key findings:
-
Tetlock (2015): Superforecasters—ordinary people who consistently outperformed intelligence analysts—achieved superior calibration through specific habits, including updating predictions as new information arrived and explicitly tracking numerical accuracy over time [1].
-
Lichtenstein & Fischhoff (1977): Demonstrated that people are systematically miscalibrated across knowledge domains, with overconfidence increasing with task difficulty; their experimental protocols established calibration curves as the standard measurement method [2].
Limitations
Limitations
Calibration alone doesn't capture everything. A forecaster who always says "50% chance" is perfectly calibrated but useless—every prediction is the same non-committal hedge. Resolution, or the ability to distinguish likely from unlikely events, matters too. The best forecasters have both calibration and resolution. Additionally, calibration assumes stable reference classes: if the world changes dramatically (like during a pandemic), past calibration may not transfer forward without recalibration. Finally, rare events pose measurement challenges—it takes hundreds of "1% risk" predictions to verify whether 1% is actually the right number.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Superforecasting: The Art and Science of PredictionPhilip Tetlock and Dan Gardner - 2015
- [2] Do Those Who Know More Also Know More About How Much They Know?Sarah Lichtenstein and Baruch Fischhoff - 1977
- [3] Calibration of Probabilities: The State of the Art to 1980Sarah Lichtenstein, Baruch Fischhoff, and David Phillips - 1982
Try it
Check your understanding
Two managers present project proposals. Manager A says there's a "90% chance of success" and succeeds in 9 of 10 similar projects. Manager B says there's a "50% chance" and succeeds in 5 of 10 projects. Which manager is better calibrated?
Show the guide's explanation
Answer: Both are equally well-calibrated
Both managers are perfectly calibrated! Manager A's 90% predictions match their 90% success rate, and Manager B's 50% predictions match their 50% success rate. Calibration is about alignment between stated probability and actual outcomes, not about having high probabilities. Manager A has higher resolution (makes more extreme predictions that distinguish more clearly), but both are equally calibrated.
A doctor tells patients there's a "20% chance of side effects" from a medication. Across 100 patients, side effects occur in 45 cases. What does this tell you about the doctor's forecasting?
Show the guide's explanation
Answer: The doctor is underconfident
The doctor is underconfident (specifically, underestimating risk). When they say "20% chance," the actual rate is 45%—more than double their estimate. Underconfidence means the stated probability is *lower* than what actually happens. This is the opposite of overconfidence, where someone says "90% sure" but is only right 60% of the time. Here, patients aren't getting accurate risk information to make informed decisions.
Your friend always says "50% chance" for every prediction—from sports outcomes to whether it will rain. After tracking 100 predictions, outcomes happen exactly 50% of the time. Is this a good forecaster?
Show the guide's explanation
Answer: Calibrated but not useful
Your friend is perfectly calibrated (50% stated → 50% actual) but has terrible resolution. Resolution is the ability to discriminate between likely and unlikely events—to say "80%" when something is very likely and "20%" when it's unlikely. Always predicting 50% provides no information for decision-making. A good forecaster needs both calibration (accurate probabilities) and resolution (meaningful distinctions between cases). This is why meteorologists can say "70%" or "30%" rather than defaulting to "maybe."
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.