Mental model
Scoring Rules & Calibration
Learn to measure and improve the accuracy of your probabilistic forecasts by honestly assessing your own uncertainty.
Discover
A weather forecaster who is always '100% certain' it will rain or '100% certain' it won't is the most useful forecaster. Is this a fact or a myth?
Is extreme confidence always a good thing?
Let's see why honest probabilities matter.
Understand
Understand
The idea that a forecaster who is always 100% certain is the most useful is a myth. Such forecasters are often overconfident and poorly calibrated—meaning their sense of certainty doesn't match reality. Scoring rules are a way to grade probabilistic predictions, rewarding both accuracy and an honest assessment of uncertainty. Calibration measures whether your '80% confident' predictions actually happen 80% of the time. For example, if a project manager says she's 90% sure a task will be done on Friday, but similar tasks are only completed on time 60% of the time, she's poorly calibrated.
Reflect on this: When I say I'm 'pretty sure' about something, what probability would I actually assign to it?
Full explanation
Full explanation
Scoring rules provide a formal process for evaluating probabilistic judgments. The process begins with an input: a specific, numbered forecast (e.g., "There is a 70% chance of rain tomorrow"). After the event occurs (it either rains or it doesn't), a scoring rule is used to assign a score based on both the forecast and the outcome. Repeating this process over time allows you to assess your overall accuracy and calibration.
A key concept is the proper scoring rule, which is designed to give you the best possible expected score only when you state your true belief. These rules discourage hedging or exaggerating your confidence. The most common is the Brier score, which penalizes you based on how far your probability was from the actual outcome (0 for didn't happen, 1 for did happen).
Calibration is the result of this long-term tracking. A well-calibrated person knows what their internal feeling of "70% certain" actually means in the real world. If you look back at all the times you were 70% confident, you should find you were right about 70% of the time. If you were only right 50% of the time, you are overconfident.
In medicine, a well-calibrated doctor can tell a patient, "There's a 20% chance this is a serious condition." This provides crucial information for deciding on risky tests, which is far more useful than a vague "it's unlikely."
Similarly, a startup founder who says, "We have a 60% chance of hitting our revenue target this quarter," allows investors to manage their expectations and risks appropriately. This is more valuable than blind optimism, as it reflects an honest and measured understanding of the business's uncertainty.
Research
Research
Scoring rules provide a formal mechanism to evaluate and improve probabilistic judgment, moving it from a vague art to a measurable skill. Research pioneered by forecasters and psychologists demonstrates that systematic tracking and feedback with proper scoring rules, like the Brier score, significantly improves a person's calibration and accuracy over time.
- Tetlock & Gardner (2015) used Brier scores to identify "Superforecasters," finding that frequent feedback with these scores was a key practice that enabled them to consistently outperform intelligence analysts with access to classified information. [1]
- Gneiting & Raftery (2007) provided a foundational review of scoring rules, distinguishing between calibration (the correspondence between forecasts and reality) and sharpness (the concentration of the predictive distributions), arguing that a good forecast must have both. [2]
- Poses et al. (1985) found that even experienced physicians were often poorly calibrated, showing systematic overconfidence in their diagnostic probability estimates, but that this could be improved with formal feedback and training. [3]
Limitations
Limitations
Scoring rules and calibration are powerful but not universally applicable. Their primary limitation is that they require clear, specific, and verifiable outcomes. They are difficult to apply to vague predictions ("the economy will improve") or forecasts so far in the future that feedback is impractical.
Furthermore, good calibration alone is insufficient. One could be perfectly calibrated by always assigning the base rate (e.g., a 50% chance for a coin flip), but this forecast is not very informative or 'sharp'. The goal is to be both well-calibrated and make confident (sharp) predictions when justified by evidence.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Superforecasting: The Art and Science of PredictionPhilip E. Tetlock & Dan Gardner - 2015
- [2] Strictly Proper Scoring Rules, Prediction, and EstimationTilmann Gneiting & Adrian E. Raftery - 2007
- [3] The accuracy of experienced physicians' probability estimates for patients with sore throats: implications for clinical decision makingR M Poses, R D Cebul, M Collins, S S Fager - 1985
- [4] How is the Metaculus Prediction scored?Metaculus - 2023
- [5] A talk by Julia Galef about calibrationJulia Galef - 2013
Try it
Check your understanding
A political analyst made 10 predictions, each with '80% confidence.' In the end, only 4 of the 10 events actually happened. What does this primarily demonstrate?
Show the guide's explanation
Answer: Overconfidence (poor calibration)
The analyst's 80% confidence was only met with a 40% success rate. This is a classic sign of overconfidence, which is a form of poor calibration. A well-calibrated forecaster's 80% predictions would be correct about 8 out of 10 times.
You are a product manager. Why is it more useful to say 'I'm 70% confident we'll ship this feature by Friday' than 'I'm pretty sure we'll ship by Friday'?
Show the guide's explanation
Answer: It creates a shared, specific understanding of risk that the team can plan around.
Assigning a number forces clarity and creates a shared reality. A 70% confidence level implies a 30% risk of not shipping, which allows the team and stakeholders to prepare a contingency plan, unlike the vague phrase 'pretty sure' which can be interpreted differently by everyone.
Which statement best reflects the goal of using a 'proper' scoring rule like the Brier score?
Show the guide's explanation
Answer: To incentivize forecasters to honestly report their true level of uncertainty.
Proper scoring rules are mathematically designed so that you achieve the best long-term average score by reporting your genuine probabilistic belief. This discourages both timid hedging and overconfident posturing, rewarding honesty above all.
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.