Mental model

Representativeness & Base-Rate Neglect

We judge how likely something is by how closely it matches a familiar example, often ignoring the actual statistical likelihood.

Discover

A city has 85 green cabs and 15 blue cabs. A witness says a hit-and-run cab was blue. The witness is 80% accurate. What's the chance the cab was actually blue?

Choose your answer

See why most people get this wrong.

Understand

Understand

We often judge how likely something is by how closely it matches a familiar example, ignoring the actual odds. In the cab problem, most people focus on the witness accuracy (80%) but overlook that green cabs are far more common (85% vs 15%). The real chance is only about 41%, not 80%. This pattern appears everywhere: doctors overestimate cancer risk from positive tests, investors chase 'hot streaks' that look like past winners, and people assume someone quiet must be a librarian not a banker despite more bankers existing. Check this: When you hear a specific story or vivid detail, ask yourself what the underlying statistics say first.

Full explanation

Full explanation

How representativeness overrides base rates

Our brains love patterns. When something matches a familiar prototype, we feel it's more likely than statistics would suggest. This mental shortcut—called the representativeness heuristic—works well in many everyday situations, but it fails systematically when base rates conflict with stereotypes. In classic studies, people consistently ignored population statistics when given a personality sketch that seemed representative of one group, even when that group was numerically rare.

Medical decisions: false positives that feel real

Doctors and patients routinely overestimate disease risk from positive test results. Consider a mammogram. If breast cancer affects 0.8% of women and a test has 90% sensitivity and 7% false positive rate, then only a small fraction of positive results indicate actual cancer. Studies have found that doctors frequently overestimate this probability. Similar errors occur with HIV tests, drug screening, and forensic evidence.

Beyond the lab: everyday pattern-matching errors

At work, representativeness makes us expect job candidates to fit familiar archetypes rather than considering what's statistically common. Investors chase funds with recent performance that matches past winners, neglecting that most hot streaks are random variation. In legal settings, jurors give too much weight to eyewitness confidence without considering how often eyewitnesses are wrong in similar situations. Political judgments suffer too: we assume a leader's personality explains events when structural factors matter more.

When does this bias strengthen or weaken?

Base-rate neglect is strongest when individuating information feels vivid and specific. Abstract statistics fade next to concrete stories. The bias weakens when base rates are presented as frequencies rather than percentages (10 out of 1000 rather than 1%), when people learn from direct experience rather than verbal description, and when decision stakes are high enough to motivate careful reasoning. Individual differences matter too: people with stronger statistical training show less neglect, though even experts can fall prey under time pressure.

Research

Research

Research on representativeness and base-rate neglect spans five decades, beginning with Kahneman and Tversky's foundational work on heuristics and biases. Their studies showed systematic departures from normative probability judgment when specific, individuating information was available. The cab problem, first described by Tversky and Kahneman, became a canonical demonstration: participants given base rates (85% green, 15% blue) and witness reliability (80% accurate) produced median estimates of 80% rather than the Bayesian posterior of roughly 41%.

  • Kahneman and Tversky (1972): Introduced representativeness as a heuristic for judging probability, showing that subjective probability depends on similarity to a prototype rather than statistical frequency, with sample size having little effect on likelihood judgments [1]
  • Tversky and Kahneman (1973): Demonstrated base-rate neglect in the lawyer-engineer problem, where participants' predictions about a person's occupation were dominated by a personality sketch rather than prior probabilities, even when base rates were explicitly provided [2]
  • Gigerenzer and Hoffrage (1995): Showed that natural frequencies (10 out of 1000) dramatically improve Bayesian reasoning compared to conditional probabilities, suggesting base-rate neglect partly reflects information format rather than fundamental irrationality [3]
  • Stengård et al. (2022): Found individual differences form a bimodal distribution—some participants nearly ignore base rates while others nearly perfectly account for them, with the former group fit by linear-additive models and the latter by Bayesian models [4]

Limitations

Limitations

Base-rate neglect is not universal or inevitable. Gigerenzer and colleagues argue that people can reason effectively about probabilities when information is presented as learned frequencies rather than single-event probabilities. Koehler (1996) documented methodological challenges: many studies used ambiguous wording, and base-rate neglect diminishes when problems are clarified and when participants experience random sampling directly. Cosmides and Tooby (1996) found dramatically better performance when information was presented in frequency formats that may match human evolutionary experience. There is also ongoing debate about whether base-rate neglect reflects cognitive limitation versus adaptive strategy—ignoring base rates can be rational when they seem irrelevant to the specific case. Individual differences matter substantially, with statistical training reducing but not eliminating the effect.

Try it

Synthesize

Choose a pattern from the guide, then pick an action to try with it.

Which pattern stands out?

What will you try?

Choose a pattern above to select an action.

Sources

Sources

Try it

Check your understanding

A city has 1,000 taxis: 850 green, 150 blue. A witness identifies a taxi as blue and is 80% accurate. What is the approximate probability the taxi was actually blue?

Show the guide's explanation

Answer: 41%

The Bayesian answer is about 41%. Start with 1000 cabs: 850 green, 150 blue. The witness correctly identifies 80% of blue cabs (120) but also misidentifies 20% of green cabs as blue (170). So 290 cabs are identified as blue, but only 120 actually are—roughly 41%. Most people say 80%, neglecting the base rate that green cabs are far more common.

Which research finding is supported by the Gigerenzer and Hoffrage studies on base-rate neglect?

Show the guide's explanation

Answer: Natural frequencies improve Bayesian reasoning compared to conditional probabilities

Gigerenzer and Hoffrage (1995) demonstrated that presenting information as natural frequencies (10 out of 1000) dramatically improves Bayesian reasoning compared to conditional probabilities (1%). This suggests base-rate neglect partly reflects information format rather than fundamental irrationality—people are better intuitive statisticians with ecologically valid representations.

In the classic cab problem, why do most people estimate 80% instead of the Bayesian answer?

Show the guide's explanation

Answer: Representativeness makes witness accuracy feel more relevant than base rates

The representativeness heuristic causes people to focus on the vivid, specific information (witness said blue, 80% accurate) while neglecting abstract base rates (15% of cabs are blue). The individuating information feels more diagnostic of the specific case, leading people to anchor on witness accuracy rather than properly integrating both pieces of information.

Keep exploring

Find another idea for the decision in front of you.

The complete Reframo library is free to read. Explore another guide whenever you are ready.