Base rate neglect
Also known as base rate fallacy or base rate bias
Base rate neglect is judging how likely something is from the details of one case while giving too little weight to how common that kind of thing is in the first place. Its status is contested: the arithmetic that makes it an error isn’t in doubt, and people usually underweight base rates in classic word problems, but how strongly they do so, and how far that carries beyond those problems, is disputed (see Evidence).
A base rate is how often something occurs in the group a case comes from: 1 in 100 people have a condition, or most cabs in a city belong to one company. The flaw is that a case’s details tell you how well it fits each possibility, not how many cases of each kind there are to begin with. A detail that fits a rare possibility well can still turn up more often under a common one, simply because the common one has far more members.
Examples
The positive test
A screening test is for a condition that affects 1 in 100 people. It detects 90% of people who have the condition, but also wrongly flags 9% of people who don’t. Your result comes back positive, and you figure there’s about a 90% chance you have it.
Picture 1,000 people taking the test. About 10 have the condition, and 9 of them test positive. The other 990 don’t have it, but 9% of them, about 89 people, test positive too. So of roughly 98 positive results, only 9 belong to people who have the condition: about 9%, not 90%. The 90% is how often the test catches the condition when it’s there. The chance you have it given a positive result also depends on how many healthy people are being tested, and here they outnumber the sick 99 to 1.
The quiet book lover
A new neighbor is quiet, tidy and never without a book. “She’s got to be a librarian,” a friend says. “Look at her. She’s exactly the type.”
The description may well fit librarians better than salespeople. But salespeople vastly outnumber librarians, so even if only a small share of them are quiet, tidy readers, there can easily be more quiet, tidy readers selling things than working in libraries. How well someone matches a type says nothing about how many people belong to that type. (Daniel Kahneman’s version of this puzzle uses farmers: in the United States, he notes, there are more than 20 male farmers for every male librarian.)
A description that says nothing
Of the 40 people a company hired this year, 30 went into sales and 10 into engineering. You’re told that one of them, Sam, “is in his thirties, well liked by his colleagues, and expected to do well.” Asked whether Sam is more likely to be in sales or engineering, you say it’s a coin flip.
Nothing in the description favors either job, so the right answer is the one you’d give with no description at all: 30 out of 40, a 75% chance of sales. The details are worthless as evidence, yet the answer treats the hiring numbers as beside the point. This is the subtlest version: the base rate isn’t outweighed by strong evidence, it’s displaced by no evidence. It mirrors a result from the original 1973 study, though later attempts to reproduce that particular result were mixed (see Evidence).
The café that can’t fail
“Their café is a sure thing. The owners have both managed restaurants, the street is busy, and the coffee is the best I’ve had.”
No statistic is mentioned, which is exactly the problem. Each detail fits success, but plenty of new cafés with experienced owners, busy streets and good coffee still close. The question the speaker skipped is how often new cafés like this one survive, and only then how much these particular strengths should raise that figure. When the base rate isn’t handed to you, neglecting it looks like confidence rather than a mistake.
Variants
- Ignoring a stated base rate: the numbers are given but the case evidence is used alone, as in the positive test above. The classic laboratory form.
- Judging by resemblance: the case is assigned to whichever group it most resembles, as with the quiet book lover. This is the same shortcut behind the Conjunction fallacy.
- Forgetting an unstated base rate: predicting from a case’s own details without asking how similar cases usually turn out, as with the café. In forecasting time and cost, this is the Planning fallacy.
- The prosecutor’s fallacy, a courtroom relative: treating the chance that an innocent person would match the evidence as if it were the chance that a matching person is innocent. Like the positive test, it leaves out how many innocent people there are who could have matched.
When it isn’t an error
- When the case evidence is highly reliable. A very accurate test or a distinctive, well-verified detail can outweigh even a low base rate. The base rate still counts, but strong enough evidence should move you most of the way.
- When no base rate really fits the case. Every case belongs to many groups (all cabs in the city, cabs out at night, cabs in that neighborhood), and they can have different rates. Setting aside a broad population figure for one that better matches the case isn’t neglect; it’s choosing the right comparison.
- When the base rate is unknown or unreliable. A made-up, outdated or unrepresentative figure deserves less weight, whichever way it points.
- When the question is about resemblance, not probability. “She seems like a librarian” is a fair description of an impression. The error is turning it into “she’s probably a librarian.”
The test: before the details, how common was each possibility, and do the details favor one by enough to overcome that?
Looks like it, but isn’t
A second, independent test
After a positive result on the screening test above, a doctor orders a different, more specific test, one that detects 90% of real cases but flags only 1% of people who don’t have the condition. It also comes back positive, and the doctor says the condition is now very likely.
Someone could object that the condition is rare, so the doctor is ignoring the base rate. But the doctor is using it, and updating it. The group this patient now belongs to is people who screened positive, and in that group about 9 in 98 have the condition, not 1 in 100. Of those 98, about 8 of the 9 who have the condition would test positive again, and only about 1 of the 89 who don’t. So roughly 8 of 9 second positives are real. (That assumes the two tests don’t make the same mistakes on the same people, which is part of why the second test is a different one.) This is the highly reliable evidence condition: a strong test applied to a better-chosen group.
An accurate test for a common condition
At a clinic, about 3 in 10 patients who come in with a particular set of symptoms turn out to have a certain infection. A test detects 99% of cases and wrongly flags 1% of people without it. A patient with those symptoms tests positive, and the doctor tells them they almost certainly have it.
This looks like trusting the test’s accuracy and nothing else, but here the base rate and the test agree. Out of 1,000 such patients, about 300 have the infection and 297 of them test positive; of the 700 who don’t, about 7 test positive. So roughly 297 of 304 positives are real, about 98%. The relevant base rate is for patients with these symptoms, not the general public, and it’s high; the test is also very accurate. Base rate neglect needs a conflict between the two, and there isn’t one.
Why it happens
Asked how likely a case is to belong to a group, people tend to answer an easier question instead: how well it fits. Kahneman and Tversky called this judging by representativeness, a heuristic (a mental shortcut). Resemblance doesn’t depend on how many members a group has, so a judgment built on resemblance leaves the base rate out. In Kahneman’s account of the “Tom W” problem, the rankings that psychology graduate students with statistical training gave fields of study by probability did not differ from rankings by similarity to the description.
Two other features make case details feel more informative than they are:
- Specific information feels more relevant than general information. A description, a witness or a test result is about this case. A base rate is about a population, and can feel like it’s about someone else. Maya Bar-Hillel proposed that people let the information that seems most specific to the case dominate. Relatedly, base rates that suggest a cause get used much more readily than bare counts. In a well-known taxicab problem, “85% of the cabs in the city are Green” tends to be ignored, while “Green cabs are involved in 85% of accidents”, mathematically the same in the problem, is given considerable weight. Kahneman’s explanation is that the second version suggests a cause, reckless Green drivers, and causes feel like facts about the individual case.
- A coherent story feels like strong evidence. In Thinking, Fast and Slow, Kahneman connects this to “what you see is all there is”: the story built from the description in front of you makes its evidence seem more telling than it is, even when you’ve been told the description might not be accurate.
Presentation matters too. Percentages like “detects 90%” and “flags 9%” each describe a different group, and combining them in your head is hard. Counts of people put everyone in the same picture, which is why the counting method below helps.
How to respond
Count people instead of combining percentages. Pick a round number of cases, say 1,000, and sort them. For the positive test above:
| Test positive | Test negative | Total | |
|---|---|---|---|
| Have the condition | 9 | 1 | 10 |
| Don’t have it | 89 | 901 | 990 |
| Total | 98 | 902 | 1,000 |
Then read down the column you actually care about: of 98 positive results, 9 belong to people with the condition. The base rate is built into the table (it’s why the bottom row is so much bigger than the top), so you can’t leave it out by accident.
This has been tested, and it helps substantially without solving the problem. In Gigerenzer and Hoffrage’s (1995) studies, 16% of answers to problems stated as probabilities were reached by a correct (Bayesian) strategy, against 46% for the same problems stated as frequencies. A 2017 meta-analysis of 35 articles found average accuracy of roughly 4% with probabilities and 24% with these “natural frequencies”. Simplified presentations and visual aids also improved performance. Counting makes the problem easier, but most people still need to set it up deliberately.
Other things to check, in yourself or in conversation:
- Ask “out of how many?” Instead of “you’re ignoring the base rate”, try “out of everyone who’d get this result, how many actually have it?” It points at the same gap without a lecture.
- When no base rate is given, find one. Ask how cases like this usually turn out before weighing what’s special about this one. Kahneman’s advice is to anchor on a plausible base rate, then question how much the evidence really tells you.
- Ask whether the details would look different under the other possibility. A description that would fit a salesperson just as well as a librarian shouldn’t move you at all.
- Don’t overcorrect. The base rate is a starting point, not the answer. Strong evidence should move you, as it did in the look-alikes above. The same principle is at work in Affirming the consequent: a result that fits a cause is evidence for it, not proof.
Evidence
Status: contested. That people underweight base rates in probability word problems has been found repeatedly for 50 years, in laypeople and in doctors. What’s disputed is how large that effect is, how far it generalizes beyond such problems, and whether “neglect” describes it. This site labels an effect robust when it has a preregistered or multi-lab replication, or a meta-analysis that survives correction for publication bias. For base rate neglect itself, neither was found in reviewing this entry: the main meta-analysis tested a remedy (presenting counts), and the largest preregistered study used that remedy’s format and found a different error. Meanwhile a prominent review argues that the effect’s size and generality have been overstated. Where the evidence falls between labels, the site chooses the less confident one.
The original studies.
- Kahneman and Tversky (1973) told participants that personality descriptions had been drawn from a group of engineers and lawyers: 70 engineers and 30 lawyers for some, the reverse for others. Both groups gave very similar probabilities that a given description belonged to an engineer. With no description, they used the base rate correctly. With a deliberately uninformative description (a 30-year-old man, married, “well liked by his colleagues”), they answered about 50% whichever group they’d been told about. The authors described base rates as “largely ignored”; a small base-rate effect was statistically significant.
- The cab problem. Participants learn that 85% of a city’s cabs belong to one company and 15% to another, and that a witness who identifies colors correctly 80% of the time says the cab in a hit-and-run came from the smaller company. The correct answer is about 41%. Bar-Hillel credits the problem to a 1972 report by Kahneman and Tversky, and Tversky and Kahneman (1980) reported that a version whose base rate invites a causal reading (the two companies have equal numbers of cabs, but one company’s cabs “are involved in 85% of accidents”) was given far more weight than the purely statistical one. In Bar-Hillel’s (1980) version, the median answer among 52 participants was 80%, the witness’s accuracy alone; only about 10% gave answers close to 41%. She reports the same pattern across many variations of the problem.
- Doctors. Casscells, Schoenberger and Graboys (1978) asked 60 physicians, residents and students at Harvard teaching hospitals about a test for a disease with a prevalence of 1 in 1,000 and a 5% false-positive rate. About 18% gave the correct answer (roughly 2%); the most common answer was 95%. Manrai and colleagues (2014) asked the same question of 61 physicians, residents and students at a Boston hospital and found nearly the same: 23% correct, with 95% again the most common answer.
Challenges and qualifications.
- Koehler (1996), a target article in Behavioral and Brain Sciences, argued that “we have been oversold” on the base rate fallacy. Across eight lawyer-engineer experiments with informative descriptions, base rates shifted judgments in every study, by 2 to 30 percentage points. With uninformative descriptions like the one in the 1973 study, results split: two studies, including the original, found base rates had little effect, while several others found strong base-rate effects. He concluded that base rates are rarely ignored outright, that their use depends on how the task is structured and presented, and that they are used more when they’re learned from experience, reliable, or relatively diagnostic. He also argued that many real situations don’t map cleanly onto the textbook Bayesian standard, because the right reference class and the reliability of a base rate are often unclear.
- Presentation format. Gigerenzer and Hoffrage (1995) and a meta-analysis by McDowell and Jacobs (2017) found that stating the information as counts substantially improves performance (see How to respond), which supports the view that the size of the effect depends heavily on how the problem is presented. Even with counts, most participants in the meta-analysis’s studies still answered incorrectly.
- Individual differences. Stengård, Juslin, Hahn and van den Berg (2022) noted that most evidence comes from problems with very low base rates and highly accurate tests. Testing a much wider range of problems, they found the pattern generalized, but participants split into two groups: some ignored the base rate almost entirely, and others took it almost fully into account.
- A reversed error. In three preregistered online studies with 2,238 participants, Pighin, Filimon and Tentori (2024) gave people Bayesian problems stated as natural frequencies. Only 14–17% answered correctly, but the most common identifiable wrong answer was the base rate alone, the opposite of neglect. The authors caution that this may not carry over to problems stated as probabilities.
Kahneman himself writes that he and Tversky originally believed base rates would always be neglected when information about the individual case was available, “but that conclusion was too strong.” What remains uncertain is how much people underweight base rates in everyday judgments, where the base rates are usually learned rather than stated, and whether one mechanism explains the pattern.
Sources
- Daniel Kahneman and Amos Tversky (1973). On the psychology of prediction. Psychological Review 80(4), 237–251.
- Amos Tversky and Daniel Kahneman (1980). Causal schemas in judgments under uncertainty. In M. Fishbein (ed.), Progress in Social Psychology, vol. 1, 49–72. Lawrence Erlbaum Associates.
- Maya Bar-Hillel (1980). The base-rate fallacy in probability judgments. Acta Psychologica 44(3), 211–233.
- Ward Casscells, Arno Schoenberger and Thomas B. Graboys (1978). Interpretation by physicians of clinical laboratory results. New England Journal of Medicine 299(18), 999–1001.
- Arjun K. Manrai, Gaurav Bhatia, Judith Strymish, Isaac S. Kohane and Sachin H. Jain (2014). Medicine's uncomfortable relationship with math: Calculating positive predictive value. JAMA Internal Medicine 174(6), 991–993.
- Gerd Gigerenzer and Ulrich Hoffrage (1995). How to improve Bayesian reasoning without instruction: Frequency formats. Psychological Review 102(4), 684–704.
- Jonathan J. Koehler (1996). The base rate fallacy reconsidered: Descriptive, normative, and methodological challenges. Behavioral and Brain Sciences 19(1), 1–17.
- Michelle McDowell and Perke Jacobs (2017). Meta-analysis of the effect of natural frequencies on Bayesian reasoning. Psychological Bulletin 143(12), 1273–1312.
- Elina Stengård, Peter Juslin, Ulrike Hahn and Ronald van den Berg (2022). On the generality and cognitive basis of base-rate neglect. Cognition 226, article 105160.
- Stefania Pighin, Flavia Filimon and Katya Tentori (2024). The impact of problem domain on Bayesian inferences: A systematic investigation. Memory & Cognition 52(4), 735–751.
- Daniel Kahneman (2011). Thinking, Fast and Slow (chapters 1, 14 and 16). Farrar, Straus and Giroux.
Last reviewed 2026-09-13.