Mental model
Weighted Averaging by Track Record
A method for combining predictions by giving more influence to sources with proven past accuracy, using scoring rules to measure and weight expertise.
Discover
Three colleagues give you different probabilities that your new product will launch on time. Sarah has been right on 80% of her past project predictions, Mike on 60%, and Jamie has no track record. If you average their predictions equally, you're ignoring a crucial signal. What's the smarter approach?
How would you combine their predictions?
Learn how to systematically weigh forecasts by proven performance.
Understand
Understand
Weighted averaging by track record means giving more influence to predictions from people or sources that have been accurate in the past, and less influence to those who haven't. It's like choosing a hiking guide based on who has successfully led trips before rather than picking someone at random. You measure past accuracy using scoring systems (lower scores are better), then use those scores to determine how much weight each prediction gets when combining them. For example, if one analyst consistently predicts outcomes more accurately than others, their forecast should count more toward your final decision. Try this: When you receive conflicting predictions, ask about each source's track record before averaging them.
Full explanation
Full explanation
Weighted averaging by track record works by first measuring each forecaster's past accuracy, then using those measurements to determine how much influence their current prediction receives. The process starts with a scoring rule that quantifies how close past predictions were to actual outcomes. Common scoring rules measure the difference between predicted probabilities and what really happened—smaller differences mean better accuracy. Once you have accuracy scores for each forecaster, you convert them to weights. Better forecasters get larger weights, worse forecasters get smaller weights. Finally, you multiply each forecaster's current prediction by their weight and sum the results.
This approach consistently outperforms simple averaging because it exploits stable individual differences in forecasting skill. Research from geopolitical forecasting tournaments shows that past performance predicts future accuracy, even after accounting for question difficulty and timing. The benefits appear once you have enough resolved predictions per forecaster to distinguish skill from luck—typically 20-40 or more. Before that, simpler methods work just as well. Interestingly, the number of forecasts someone makes (how often they update their beliefs) also predicts accuracy, especially early on.
The technique applies beyond formal forecasting. Investment teams can weight analysts by their stock-picking track record. Product teams can prioritize features suggested by designers whose past choices drove user engagement. Medical teams might give more weight to diagnoses from specialists with documented accuracy in similar cases. Even news consumers can mentally weight sources by their history of reliable reporting versus sensationalism.
Weighted by track record differs from extremizing, which pushes aggregated probabilities toward certainty (closer to 0% or 100%). Both methods aim to improve aggregate accuracy, but they operate differently: weighting changes who gets influence, while extremizing adjusts the final combined probability to account for overlapping information among forecasters. Sophisticated systems often use both—weighting by track record first, then extremizing the weighted aggregate.
Research
Research
Weighted averaging by track record is grounded in the finding that forecasting ability is a relatively stable trait that can be measured and leveraged for better aggregation. The Brier score, a strictly proper scoring rule developed in 1950, measures mean squared error between predicted probabilities and actual outcomes—lower scores indicate better accuracy and provide an honest incentive for truthful reporting.
- Satopää et al. (2014): Found that extremizing aggregated forecasts toward more confident probabilities improves accuracy when experts have access to different information, with optimal extremization factors derived from historical data.[1]
- Mellers et al. (2015): Demonstrated that top performers identified through past performance (Superforecasters) consistently outperformed others, and that weighted aggregation using their judgments significantly improved crowd accuracy while preserving wisdom-of-crowds benefits.[2]
- The OTexts forecasting handbook notes that while simple averaging is remarkably hard to beat, weighted approaches can provide gains when you have reliable information about individual forecaster skill.[3]
However, simpler hierarchical models performed nearly as well with much less computational cost. The key insight: most benefits come from having any valid measure of skill, not from maximizing methodological sophistication.
Limitations
Limitations
Weighted averaging by track record suffers from the cold start problem—you need past performance data to assign weights, but new forecasters have no history. Solutions include using individual differences (fluid intelligence, numerical reasoning, actively open-minded thinking) as proxies until performance data accumulates. The method also assumes past skill predicts future performance, which may not hold if domains change dramatically or if the task shifts. Extremely small sample sizes (fewer than 10-20 resolved forecasts per person) produce noisy skill estimates that may degrade rather than improve aggregation. There's also debate about whether complex IRT-based skill measurements provide practically meaningful benefits over simpler accuracy averages, especially for real-world applications where computational efficiency matters.
Try it
Synthesize
Choose a pattern from the guide, then pick an action to try with it.
Which pattern stands out?
What will you try?
Choose a pattern above to select an action.
Sources
Sources
- [1] Combining and Extremizing Real-Valued ForecastsVille Satopää, Jonathan Baron, Dean Foster, Barbara Mellers, Philip Tetlock, Lyle Ungar - 2014
- [2] Identifying and Cultivating Superforecasters as a Method of Improving Probabilistic PredictionsBarbara Mellers, Lyle Ungar, Jonathan Baron, et al. - 2015
- [3] Forecasting: Principles and Practice (2nd ed) - Forecast CombinationsRob J Hyndman, George Athanasopoulos - 2018
- [4] Brier scoreWikipedia contributors - 2024
Try it
Check your understanding
A team has three analysts predicting quarterly revenue. Alex has a Brier score of 0.15 (excellent), Taylor has 0.25 (good), and Jordan has 0.40 (fair) on past predictions. They all give different revenue forecasts. What's the first step in applying weighted averaging by track record?
Show the guide's explanation
Answer: Convert Brier scores to weights (Alex gets highest, Jordan lowest)
Weighted averaging by track record requires first converting accuracy scores into weights. Since lower Brier scores indicate better accuracy, Alex's 0.15 earns the highest weight, Taylor's 0.25 a moderate weight, and Jordan's 0.40 the lowest weight. The weighted aggregate then becomes: (Alex's forecast × high weight) + (Taylor's forecast × medium weight) + (Jordan's forecast × low weight), divided by total weights.
Why does weighted averaging by track record typically outperform simple averaging of predictions?
Show the guide's explanation
Answer: Forecasting skill is a stable trait that varies across individuals
Research from forecasting tournaments shows that past performance is a strong predictor of future accuracy—people who forecasted well before tend to forecast well again. Weighted averaging exploits this stable individual variation by giving more influence to skilled forecasters and less to less-skilled ones. Simple averaging treats all forecasters as equally good, which sacrifices accuracy when skill differences exist.
When does weighted averaging by track record provide the LEAST benefit over simple averaging?
Show the guide's explanation
Answer: With fewer than 10-20 resolved predictions per forecaster
The cold start problem limits weighted averaging's benefits early on. With fewer than 10-20 resolved forecasts per person, skill estimates are too noisy to reliably identify who's actually better. At this stage, simple averaging or weighting by proxy variables (like intelligence tests or numeracy) performs just as well. The benefits of performance-based weighting emerge once you have enough data to distinguish genuine skill from luck.
Keep exploring
Find another idea for the decision in front of you.
The complete Reframo library is free to read. Explore another guide whenever you are ready.