Regression fallacy
Also known as regressive fallacy or regression to the mean fallacy
The regression fallacy is giving the credit (or blame) for a change to something that happened in between, when an unusually high or low result was simply returning toward average. As a cognitive bias its status is contested: that extreme results tend to be followed by less extreme ones is arithmetic, and so is the damage it does to studies that ignore it, but how strongly and how generally people fail to expect it is disputed (see Evidence).
The underlying pattern is called regression to the mean. Almost any result mixes something stable (skill, health, ability) with things that vary from one occasion to the next (luck, sleep, an easy or hard day). When you pick out a result because it was extreme, you have mostly picked an occasion when the variable part happened to push the same way, and that part doesn’t carry over. The next result usually lands closer to average with no cause at all. The flaw is treating that before-and-after change as evidence that whatever happened in between made the difference, when picking an extreme starting point all but guaranteed a change in that direction.
Examples
The remedy that worked
A man’s back pain flares up so badly that he can barely sit through work. On a friend’s advice he starts wearing a magnetic wristband, and within two weeks the pain has eased. “I was skeptical, but it worked,” he tells everyone.
Pain that comes and goes is most likely to drive someone to try something new when it’s at its worst, and from its worst the likeliest direction is better. He would probably have improved over the next two weeks with or without the wristband, so his improvement can’t tell the two apart. That doesn’t show the wristband does nothing; it shows this experience isn’t evidence either way.
The cover jinx
A basketball player has the best season of her career and is featured on the cover of a national sports magazine. The next season her scoring drops noticeably. Fans start talking about a cover jinx: the attention got to her.
Her decline is real, but the cause attached to it isn’t needed. The cover went to whoever had an exceptional season, and an exceptional season is usually real skill plus unusually good luck. Next season she keeps the skill but not, on average, the luck. The jinx story invents a cause for a drop that selecting her for an extreme season made likely. She is probably still very good, just less extreme.
Praise that “backfires”
After each weekly quiz, a teacher praises the students with the highest scores and has a stern word with those who scored lowest. She notices that the praised students usually do worse the following week and the scolded ones do better. She concludes that praise makes students complacent and criticism motivates them.
Both groups were picked for extreme scores, so both would be expected to move toward the middle the next week whatever she said. Because praise always follows the high scores and criticism the low ones, regression alone makes criticism look effective and praise look harmful. This mirrors an observation Tversky and Kahneman reported from flight training, where instructors drew the same lesson about praising smooth landings and criticizing rough ones (see Evidence).
The tutoring program’s results
A school district tests every third-grader in the spring and enrolls the lowest-scoring tenth in a summer reading program. In the fall, those students are retested, and their average score has risen by 12 points. The district’s report says the program raised reading scores.
This is the research version, and it looks respectable because it has numbers and a before-and-after comparison. But the students were chosen for scoring low on one day, and some of those low scores included bad luck. Retested, the group would be expected to score higher even with no program. Without a comparison group of similarly low scorers who didn’t get the program, the 12 points mix the program’s effect (if any) with regression. Research methods texts list this as a standard threat to a study’s conclusions: statistical regression, when a group is selected for extreme scores.
Variants
- Crediting a fix: a remedy, a new policy or a new manager gets credit for improvement that followed a bad spell. Morton and Torgerson give the example of extra resources for the worst-ranked hospitals: most will climb the rankings, whatever the resources did.
- Blaming a jinx or a change: a drop after an exceptional result gets a causal story, as with the cover jinx.
- The reward-and-punishment illusion: when rewards follow good performance and punishments follow bad, regression makes punishment look like it works and reward look like it backfires.
- Regression artifacts in studies: before-and-after studies of people selected for unusually high or low measurements, without a comparison group, as in the tutoring example.
When it isn’t an error
- When a comparison group was selected the same way. If similarly extreme people who didn’t get the treatment regress too, the difference between the groups is evidence of an effect. Random assignment is the most reliable way to get such a group.
- When the starting point wasn’t picked for being extreme. A change measured from a stable baseline, such as an average of many earlier measurements, isn’t inflated by selection.
- When the change goes beyond the average and stays there. Regression pulls results toward the typical level; it doesn’t push a whole group well past it and keep it there over repeated measurements.
- When the measurement barely varies. Regression is larger the less two measurements agree, and small when a measurement is very reliable and the thing measured is stable.
- Expecting regression is not the error. Predicting that a record-breaking season will be followed by a more ordinary one is good reasoning. The error is attaching a cause to that return.
The test: was this case picked because it was extreme, and would it have moved toward average even if nothing had been done?
Looks like it, but isn’t
A drug that beats the placebo
Adults whose blood pressure was high at a screening are randomly assigned to take a new drug or a placebo. After eight weeks, the average reading in the placebo group has fallen by 5 points and in the drug group by 14. The researchers conclude the drug lowers blood pressure.
Both groups were selected for high readings, so both regress, and the placebo group shows roughly how much. The researchers don’t credit the drug with the whole 14 points, only with its advantage over a group that regressed the same way. That is the comparison group condition.
A change against a stable baseline
For a year, a bakery has thrown away between 38 and 45 unsold loaves every week. It switches to baking smaller batches through the day, and waste falls to between 15 and 20 loaves a week, where it stays for the next six months.
The bakery didn’t make the change in response to one unusually bad week; a year of records shows the normal range, and the new level is far outside it and has held. Regression can explain a bad week followed by a better one, not a lasting shift to a level the bakery never reached before. That is the stable baseline condition.
Why it happens
Tversky and Kahneman identified two separate failures. First, people don’t expect regression in many situations where it is bound to happen. Second, when they do notice it, they often invent causal explanations for it. Their account of the first is the same shortcut behind Base rate neglect, representativeness: people expect a prediction to be as extreme as the evidence it’s based on, so an exceptional performance “should” be followed by another one. When it isn’t, something must have changed, and a cause is easy to find.
Everyday life also lines up causes with extremes. People try remedies when they feel worst, praise when things went best, and bring in a new approach after a bad run. So the action and the regression arrive together, again and again, and the pattern looks like experience rather than coincidence. As Tversky and Kahneman put it, because rewards typically follow good performance and punishments poor performance, “by regression alone” behavior is most likely to improve after punishment and to deteriorate after reward.
The same pattern can mislead research on biases themselves. The Dunning–Kruger effect entry discusses regression to the mean as a live alternative explanation for its central finding.
How to respond
- Ask how the case was chosen. If it came to attention because it was unusually good or bad, expect some return toward average before looking for a cause.
- Ask what would have happened with nothing done. A comparison group, selected the same way, is the direct answer. In research, Barnett and colleagues conclude that regression’s effect can be reduced through better study design and suitable statistical methods. A control group selected the same way, ideally by random assignment, regresses just as much as the treated group.
- Measure more than once. An average of several readings is less affected by one lucky or unlucky occasion. Morton and Torgerson describe averaging repeated blood pressure measurements for this reason, and report that among women treated for osteoporosis, more than 80% of those who lost bone in the first year gained bone in the second with no change in treatment.
- Watch for the reward-and-punishment pattern. If criticism seems to work and praise seems to backfire, check whether each simply followed an extreme result.
- Don’t expect knowing the statistics to be enough. In Kahneman and Tversky’s 1973 study, most of the psychology graduate students tested, who would be expected to know the statistics, still gave non-regressive answers (see Evidence).
Evidence
Status: contested. Regression to the mean is a mathematical property of imperfectly related measurements, and its power to mislead uncontrolled studies is a textbook point of research design, not an open question. What this label covers is the psychological claim: that people make predictions that are not regressive enough and explain regression causally. That has been reported repeatedly since 1973, but in reviewing this entry no published preregistered or multi-lab replication, or meta-analysis, of those original findings was found, and later work disputes how general the pattern is. Where evidence falls between labels, the site chooses the less confident one.
The statistical phenomenon.
- Galton (1886) documented it by comparing the heights of adult children with the average height of their parents: children of taller-than-average parents tended to be shorter than their parents, and children of shorter-than-average parents taller. He called it “regression towards mediocrity.”
- Campbell and Stanley, in a guide to research design first published in 1963, listed statistical regression among the threats to a study’s internal validity (whether its result reflects the treatment): when groups are selected for extreme scores, their retest scores tend to move toward the mean.
- Barnett, van der Pols and Dobson (2005) describe regression to the mean as something that “can make natural variation in repeated data look like real change,” more noticeable with greater measurement error and when follow-up is examined only in a subgroup selected on its baseline value. They conclude it “should always be considered as a possible cause of an observed change.” Morton and Torgerson (2003) survey its effects on clinical, public health and management decisions.
The original psychological studies.
- Kahneman and Tversky (1973) argued that intuitive predictions follow representativeness and therefore fail to regress. As Tversky and Kahneman (1974) summarize one study, participants read descriptions of student teachers’ practice lessons; some rated the lesson’s quality as a percentile, others predicted each teacher’s standing five years later. The two sets of judgments were identical: the prediction of a remote outcome was as extreme as the evaluation it was based on. In another study, as summarized by the team now replicating it, 108 psychology graduate students were told that a person scored 140 on an IQ test and asked for a 95% range for the person’s true score: 73 gave a range centered on 140 (no regression), 24 a range shifted toward the mean, and 11 one shifted away from it.
- The flight instructors. Tversky and Kahneman (1974), citing their 1973 paper, report, as an observation from a discussion of flight training rather than an experiment, that experienced instructors found praise for an exceptionally smooth landing was typically followed by a poorer one and harsh criticism after a rough landing by an improvement, and concluded that rewards hurt learning and punishments help.
Qualifications and challenges.
- Ganzach and Krantz (1990, 1991) found that intuitive numerical predictions “can be somewhat regressive,” more so at low than at high values of the predictor, and became more moderate after experience with tasks where several factors determine the outcome, or with feedback. They argued this moderation doesn’t come from applying an abstract rule of regression to the mean, but from other processes, such as filling in unknown factors with average values.
- Cahan and Snapiri (2008) proposed that insufficiently regressive predictions reflect a different idea of what makes a prediction good (reproducing the spread and character of real outcomes rather than minimizing error), not simply a mistaken reliance on representativeness, and reported results supporting that view.
- Fiedler and Unkelbach (2014) argue that human judgments are themselves regressive whenever they are less than perfectly tied to the evidence, and that several well-known biases can be partly explained as regression effects. This complicates the simple picture of people who never regress enough.
Replications. Chan and Feldman note that most replications of the 1973 paper have focused on its base-rate problems rather than its prediction and regression problems. Their Registered Report, a planned close replication of most of the 1973 studies including the IQ problem, was accepted in principle in 2024; no published results were found in reviewing this entry. A 2025 preprint by Harris, Schulte and Alves, reporting four preregistered experiments (601 participants in total), found that people failed to account for regression when interpreting other people’s probability estimates, but not in the way the original studies describe: instead of correcting for the regressiveness of those estimates, participants moved them even further toward the middle. It tests a related claim rather than the original one, and is a preprint.
What remains uncertain is how often people fail to anticipate regression in everyday judgments, as opposed to word problems, and whether a single mechanism explains it.
Sources
- Francis Galton (1886). Regression towards mediocrity in hereditary stature. Journal of the Anthropological Institute of Great Britain and Ireland 15, 246–263.
- Donald T. Campbell and Julian C. Stanley (1966). Experimental and quasi-experimental designs for research. Rand McNally.
- Daniel Kahneman and Amos Tversky (1973). On the psychology of prediction. Psychological Review 80(4), 237–251.
- Amos Tversky and Daniel Kahneman (1974). Judgment under uncertainty: Heuristics and biases. Science 185(4157), 1124–1131.
- Yoav Ganzach and David H. Krantz (1990). The psychology of moderate prediction. Organizational Behavior and Human Decision Processes 47(2), 177–204.
- Yoav Ganzach and David H. Krantz (1991). The psychology of moderate prediction. Organizational Behavior and Human Decision Processes 48(2), 169–192.
- Veronica Morton and David J. Torgerson (2003). Effect of regression to the mean on decision making in health care. BMJ 326(7398), 1083–1084.
- Adrian G. Barnett, Jolieke C. van der Pols and Annette J. Dobson (2005). Regression to the mean: What it is and how to deal with it. International Journal of Epidemiology 34(1), 215–220.
- Sorel Cahan and Tchia Snapiri (2008). Intuitive prediction: Ecological validity versus representativeness. Journal of Behavioral Decision Making 21(3), 297–316.
- Klaus Fiedler and Christian Unkelbach (2014). Regressive judgment. Current Directions in Psychological Science 23(5), 361–367.
- Hong Ching Chan and Gilad Feldman (2024). Representativeness heuristic in intuitive predictions: Replication Registered Report of problems reviewed in Kahneman and Tversky (1973) [Stage 1]. Stage 1 Registered Report manuscript, recommended by Peer Community in Registered Reports.
- Chris Harris, Anna Schulte and Hans Alves (2025). Regressive judgments: How people fail to account for regression in other people's judgments. PsyArXiv preprint.
Last reviewed 2026-09-13.