Mental model

Weighted Averaging by Track Record

A method for combining predictions by giving more influence to sources with proven past accuracy, using scoring rules to measure and weight expertise.

Discover

Three colleagues give you different probabilities that your new product will launch on time. Sarah has been right on 80% of her past project predictions, Mike on 60%, and Jamie has no track record. If you average their predictions equally, you're ignoring a crucial signal. What's the smarter approach?

How would you combine their predictions?

Learn how to systematically weigh forecasts by proven performance.

Understand

Understand

Weighted averaging by track record means giving more influence to predictions from people or sources that have been accurate in the past, and less influence to those who haven't. It's like choosing a hiking guide based on who has successfully led trips before rather than picking someone at random. You measure past accuracy using scoring systems (lower scores are better), then use those scores to determine how much weight each prediction gets when combining them. For example, if one analyst consistently predicts outcomes more accurately than others, their forecast should count more toward your final decision. Try this: When you receive conflicting predictions, ask about each source's track record before averaging them.

Full explanation

Full explanation

Weighted averaging by track record works by first measuring each forecaster's past accuracy, then using those measurements to determine how much influence their current prediction receives. The process starts with a scoring rule that quantifies how close past predictions were to actual outcomes. Common scoring rules measure the difference between predicted probabilities and what really happened—smaller differences mean better accuracy. Once you have accuracy scores for each forecaster, you convert them to weights. Better forecasters get larger weights, worse forecasters get smaller weights. Finally, you multiply each forecaster's current prediction by their weight and sum the results.

This approach consistently outperforms simple averaging because it exploits stable individual differences in forecasting skill. Research from geopolitical forecasting tournaments shows that past performance predicts future accuracy, even after accounting for question difficulty and timing. The benefits appear once you have enough resolved predictions per forecaster to distinguish skill from luck—typically 20-40 or more. Before that, simpler methods work just as well. Interestingly, the number of forecasts someone makes (how often they update their beliefs) also predicts accuracy, especially early on.

The technique applies beyond formal forecasting. Investment teams can weight analysts by their stock-picking track record. Product teams can prioritize features suggested by designers whose past choices drove user engagement. Medical teams might give more weight to diagnoses from specialists with documented accuracy in similar cases. Even news consumers can mentally weight sources by their history of reliable reporting versus sensationalism.

Weighted by track record differs from extremizing, which pushes aggregated probabilities toward certainty (closer to 0% or 100%). Both methods aim to improve aggregate accuracy, but they operate differently: weighting changes who gets influence, while extremizing adjusts the final combined probability to account for overlapping information among forecasters. Sophisticated systems often use both—weighting by track record first, then extremizing the weighted aggregate.

Research

Research

Weighted averaging by track record is grounded in the finding that forecasting ability is a relatively stable trait that can be measured and leveraged for better aggregation. The Brier score, a strictly proper scoring rule developed in 1950, measures mean squared error between predicted probabilities and actual outcomes—lower scores indicate better accuracy and provide an honest incentive for truthful reporting.

  • Satopää et al. (2014): Found that extremizing aggregated forecasts toward more confident probabilities improves accuracy when experts have access to different information, with optimal extremization factors derived from historical data.[1]
  • Mellers et al. (2015): Demonstrated that top performers identified through past performance (Superforecasters) consistently outperformed others, and that weighted aggregation using their judgments significantly improved crowd accuracy while preserving wisdom-of-crowds benefits.[2]
  • The OTexts forecasting handbook notes that while simple averaging is remarkably hard to beat, weighted approaches can provide gains when you have reliable information about individual forecaster skill.[3]

However, simpler hierarchical models performed nearly as well with much less computational cost. The key insight: most benefits come from having any valid measure of skill, not from maximizing methodological sophistication.

Limitations

Limitations

Weighted averaging by track record suffers from the cold start problem—you need past performance data to assign weights, but new forecasters have no history. Solutions include using individual differences (fluid intelligence, numerical reasoning, actively open-minded thinking) as proxies until performance data accumulates. The method also assumes past skill predicts future performance, which may not hold if domains change dramatically or if the task shifts. Extremely small sample sizes (fewer than 10-20 resolved forecasts per person) produce noisy skill estimates that may degrade rather than improve aggregation. There's also debate about whether complex IRT-based skill measurements provide practically meaningful benefits over simpler accuracy averages, especially for real-world applications where computational efficiency matters.

Try it

Synthesize

Choose a pattern from the guide, then pick an action to try with it.

Which pattern stands out?

What will you try?

Choose a pattern above to select an action.

Sources

Sources

Try it

Check your understanding

A team has three analysts predicting quarterly revenue. Alex has a Brier score of 0.15 (excellent), Taylor has 0.25 (good), and Jordan has 0.40 (fair) on past predictions. They all give different revenue forecasts. What's the first step in applying weighted averaging by track record?

Show the guide's explanation

Answer: Convert Brier scores to weights (Alex gets highest, Jordan lowest)

Weighted averaging by track record requires first converting accuracy scores into weights. Since lower Brier scores indicate better accuracy, Alex's 0.15 earns the highest weight, Taylor's 0.25 a moderate weight, and Jordan's 0.40 the lowest weight. The weighted aggregate then becomes: (Alex's forecast × high weight) + (Taylor's forecast × medium weight) + (Jordan's forecast × low weight), divided by total weights.

Why does weighted averaging by track record typically outperform simple averaging of predictions?

Show the guide's explanation

Answer: Forecasting skill is a stable trait that varies across individuals

Research from forecasting tournaments shows that past performance is a strong predictor of future accuracy—people who forecasted well before tend to forecast well again. Weighted averaging exploits this stable individual variation by giving more influence to skilled forecasters and less to less-skilled ones. Simple averaging treats all forecasters as equally good, which sacrifices accuracy when skill differences exist.

When does weighted averaging by track record provide the LEAST benefit over simple averaging?

Show the guide's explanation

Answer: With fewer than 10-20 resolved predictions per forecaster

The cold start problem limits weighted averaging's benefits early on. With fewer than 10-20 resolved forecasts per person, skill estimates are too noisy to reliably identify who's actually better. At this stage, simple averaging or weighting by proxy variables (like intelligence tests or numeracy) performs just as well. The benefits of performance-based weighting emerge once you have enough data to distinguish genuine skill from luck.

Keep exploring

Find another idea for the decision in front of you.

The complete Reframo library is free to read. Explore another guide whenever you are ready.

Weighted Averaging by Track Record | Reframo