Mental model

Brier Score

A mathematical tool that rewards confidence when you're right and penalizes you harshly when you're wrong—exposing the difference between feeling certain and being correct.

Discover

Two people each predict a 70% chance of rain. One checks weather data and satellite patterns; the other just feels optimistic in their gut. Who gets the better score?

Whose prediction scores better?

Let's see how Brier scores really work.

Understand

Understand

A Brier score measures how good your probability predictions are by comparing what you said would happen to what actually happened. Think of it like a truthfulness meter for predictions: if you say there's a 90% chance of something and you're right, you get rewarded—but if you're wrong, you get penalized much more heavily than if you'd only said 60%. This is why two people can both predict 70% chance of rain and get identical scores, even if one used deep research and the other just guessed—the score only measures the final probability you picked, not how carefully you thought about it. Try this: Remember two different people can pick the same number and get identical scores—one might research for hours, another might guess randomly. The math only sees the final probability.

Full explanation

Full explanation

Brier scores work by taking your predicted probability, squaring the difference from what actually happened (1 for yes, 0 for no), and averaging across all your predictions. A perfect score is 0—every prediction matched reality exactly. A terrible score is 1—you were maximally confident and completely wrong every time. Most real-world forecasts fall somewhere between.

If you say 95% and you're wrong, your penalty is (0.95 - 0)² = 0.9025. But if you'd humbly said 55% and were wrong, your penalty would only be (0.55 - 0)² = 0.3025. This design deliberately rewards calibrated uncertainty—being appropriately confident when you have good evidence and appropriately cautious when you don't.

Consider a medical diagnosis scenario: a doctor predicts 80% chance that a patient has strep throat. If the test comes back positive, the doctor's Brier score for that prediction is (0.80 - 1)² = 0.04. If the test comes back negative, the penalty is (0.80 - 0)² = 0.64. Now imagine the doctor had said 40% instead—still favoring strep but with more uncertainty. A wrong diagnosis now costs only (0.40 - 0)² = 0.16. The system doesn't care that the doctor ordered extra tests; it only cares that the final probability was too confident given the outcome.

This reveals a key insight: Brier scores measure calibration, not information-gathering. Two pundits can both predict 65% chance that a candidate wins an election, and they'll receive identical scores regardless of whether one spent months analyzing polling data while the other flipped a coin. This isn't a flaw—it's exactly what makes Brier scores useful for comparing forecasting performance across different domains and methods.

Research

Research

Brier scores, introduced by Glenn Wilson Brier in 1950, measure the mean squared error of probabilistic predictions and are widely used for evaluating forecast calibration across fields from meteorology to machine learning.

  • Brier (1950): Introduced the scoring rule to evaluate weather forecasters [1].

Gloss: Calibration means your stated probabilities match your observed frequencies over time (you're right 80% of the time when you say 80%). Proper scoring rule means you can't improve your expected score by predicting anything other than your true beliefs.

Limitations

Limitations

Brier scores treat all errors symmetrically—being overconfident when wrong is penalized the same as being underconfident when right, which may not reflect real-world consequences in asymmetric domains (like medical diagnosis where false negatives cost more than false positives). Additionally, Brier scores only evaluate single-event predictions, not full probability distributions over multiple outcomes.

Try it

Synthesize

Choose a pattern from the guide, then pick an action to try with it.

Which pattern stands out?

What will you try?

Choose a pattern above to select an action.

Try it

Check your understanding

A forecaster predicts 80% chance of victory for Candidate A. If Candidate A loses, what is the Brier score contribution for this single prediction?

Show the guide's explanation

Answer: 0.64

The Brier score contribution is (predicted - actual)² = (0.80 - 0)² = 0.64. The large penalty reflects the cost of high confidence when wrong—this is exactly how Brier scores discourage overconfident predictions.

Two friends both predict 60% chance that their favorite team wins a game. Alex spends hours analyzing player stats; Maya picks 60% at random. They receive identical Brier scores. What does this demonstrate?

Show the guide's explanation

Answer: Brier scores measure calibration, not effort or methodology

Brier scores evaluate the final probability you chose against what actually happened—they don't reward how you got there. This is a feature, not a bug: it lets us compare forecasting accuracy across different methods and domains.

You're evaluating two meteorologists. Forecaster A has a Brier score of 0.08; Forecaster B has 0.22. Both predicted rain for tomorrow with 70% probability. Who should you trust more and why?

Show the guide's explanation

Answer: Forecaster A—lower Brier scores indicate better historical calibration

Brier scores are like golf: lower is better. 0.08 means Forecaster A's predictions have historically been very close to actual outcomes, while 0.22 suggests Forecaster B's track record includes more confident misses or poorly calibrated probabilities.

Keep exploring

Find another idea for the decision in front of you.

The complete Reframo library is free to read. Explore another guide whenever you are ready.

Brier Score | Reframo