Principle
Regression to the mean
Also known as regression toward the mean, regression towards the mean or statistical regression
Regression to the mean is the tendency for a case that was picked out because it was extreme on one measurement to be less extreme on the next. The student with the top score on one test usually scores a little lower on the next; the team with the worst month usually has a better one after it. Nothing has to happen in between.
This is a statistical principle, a consequence of how imperfectly related measurements behave. It isn’t a force that pulls things toward average, it isn’t a psychological effect measured in experiments, and it isn’t an error. The error is explaining it: giving the credit or blame for the return to some cause that happened in between, which this site covers as the regression fallacy.
Example
A bowling league posts its weekly scores. The ten bowlers with the highest scores in week one are recognized at the bar afterward. In week two, eight of the ten score lower than they did in week one. Meanwhile, most of the ten lowest scorers from week one do better in week two.
Nobody’s skill changed in a week. A bowler’s score mixes something stable (how good they are) with things that vary from night to night (a lucky split, a sore wrist, a sticky lane). The top ten of week one were picked for having high scores, so as a group they include bowlers who are good and had a good night. Their skill comes back in week two; their good night, on average, doesn’t. The same reasoning, in reverse, lifts the bottom ten. Most of the top ten are probably still above average in week two, just less far above it.
Why it happens
Almost any measurement combines a stable part and a part that varies from one occasion to the next: measurement error, luck, a good or bad day. When you select cases for being extreme, you select two things at once: cases whose stable part is extreme, and cases where the varying part happened to push the same way. Only the first carries over to a second measurement. So the group’s second average lands between its first average and the overall average.
Bland and Altman state the general condition: regression toward the mean occurs “whenever we select an extreme group based on one variable and then measure another variable for that group.” The second variable can be the same thing measured later (weight next month), a related thing (a child’s height against a parent’s), or a different way of measuring it.
How much
The size of the effect depends on how closely the two measurements agree. In the simplest case, where both have the same average and spread, Bland and Altman give a rule of thumb: the group’s second average will be about r times as far from the overall average as its first, where r is the correlation between the two measurements (a number from 0, no relationship, to 1, perfect agreement).
| Correlation between the two measurements | A group selected 20 points above average is, on the second measurement, about |
|---|---|
| 1 (perfect) | 20 points above: no regression |
| 0.8 | 16 points above |
| 0.5 | 10 points above |
| 0 (none) | at the average: complete regression |
So regression is large when a measurement is noisy or the thing measured fluctuates, and small when a measurement is reliable and stable. Barnett, van der Pols and Dobson note that it becomes more noticeable with increasing measurement error, and when follow-up measurements are examined only in a subgroup selected on its first value. Bland and Altman add that unless two measurements are perfectly correlated, regression toward the mean occurs, “so it always occurs in practice.”
A near-perfect correlation shows the other end. When Horace Secrist tracked average July temperatures in 191 American cities, he found no regression at all, as Stephen Stigler recounts. With climates that varied so much from city to city, a city’s temperature one year predicted the next almost exactly.
It runs in both directions
Regression isn’t something that happens over time. It happens whenever you select on one measurement and look at another, in either order.
Francis Galton, who named it, compared adult children’s heights with the average of their two parents’ heights. As Bland and Altman describe his data, children and parents had the same average height, 68.2 inches. Parents averaging between 70 and 71 inches had children averaging 69.5 inches: closer to the average. But starting from the children gives the same pattern: children between 70 and 71 inches had parents averaging 69.0 inches. Tall children have, on average, less tall parents, just as tall parents have less tall children. Bland and Altman: “This is a statistical, not a genetic phenomenon.”
This is also why regression doesn’t make everyone average. If it did, it couldn’t work backward as well as forward. Some of the cases that weren’t extreme the first time are extreme the second time, so the population can stay as spread out as it was while each group selected for being extreme drifts toward the middle. Regression describes the average of a group chosen for being extreme, not a squeeze on the whole population, and not a guarantee about any one individual.
A book-length mistake
Secrist, a professor of statistics at Northwestern University, published The Triumph of Mediocrity in Business in 1933. As Stigler recounts it, the book ran to 468 pages, with 140 tables and 103 charts. Secrist grouped 49 department stores by their 1920 profits and followed each group’s average over the decade: the most profitable stores declined toward the middle and the least profitable rose toward it. He found the same pattern in 73 series of business figures, from groceries to railroads and banks, and concluded that competition drives businesses toward mediocrity. Several reviewers praised the book.
Harold Hotelling’s review in the Journal of the American Statistical Association pointed out that if the firms were grouped by their values in the last year instead of the first, “the lines would diverge. Thus from the same data one may demonstrate stability or instability according to taste. The seeming convergence is a statistical fallacy, resulting from the method of grouping.” If businesses were really converging, Hotelling noted, the spread across individual firms would narrow over time, and it didn’t. In a later exchange, Hotelling compared proving the result with years of business data to “proving the multiplication table by arranging elephants in rows and columns.”
Where it misleads
- Crediting a cause. A remedy tried at the worst moment, a coach replaced after a bad run, a policy aimed at the worst performers: each is followed by improvement that regression predicted anyway. This is the Regression fallacy, a close relative of post hoc reasoning.
- Treating high readings. Bland and Altman note that if people with high blood pressure are treated and measured again, their average will be lower, and “This should not be interpreted as showing the effect of the treatment”, because it would fall even without treatment.
- “It works best for the worst cases.” For the same reason, a treatment will seem to help most in people who started with the most extreme readings, a pattern Bland and Altman say “we would expect to observe” even in untreated patients.
- Comparing two measurements of the same thing. Bland and Altman describe studies that compared people’s self-reported weight with their measured weight and concluded that heavy people underestimate their weight and light people overestimate it. Regression predicts exactly that pattern even if self-reports are as accurate as the scale, and analyzing the same data the other way round produces the opposite conclusion.
- Research on biases. The Dunning–Kruger effect entry discusses regression to the mean as a live alternative explanation for why low scorers seem to overestimate themselves and high scorers to underestimate.
- Selecting the best. Bland and Altman point out that because reviewers judge quality with error, an editor who accepts the papers judged best will find their average quality lower than expected, and the rejected papers’ average quality higher.
What does help
- A comparison group selected the same way. If everyone was chosen for a high reading, a group that gets no treatment regresses just as much, so the difference between the groups is what’s left for the treatment to explain. Bland and Altman say separating real change from regression “is best done by using a randomised control group.”
- Better baselines. Averaging several measurements before selecting anyone, instead of relying on a single reading, reduces the part of an extreme value that was just a bad day. Morton and Torgerson describe averaging repeated blood pressure readings for this reason. Barnett and colleagues conclude that regression’s effect “can be alleviated through better study design and use of suitable statistical methods.”
- Expecting it in predictions. When the best guess about a case comes from one extreme result, predict something between that result and the case’s usual level. This is the sound reasoning that the gambler’s fallacy entry contrasts with expecting results to “even out.”
Limits
- Regression doesn’t show that nothing happened. A real improvement and regression can happen at the same time. A comparison group, not the principle alone, is what tells you how much of a change was which.
- It needs selection. A group chosen at random, or for reasons unrelated to the measurement, doesn’t regress as a group. The question to ask is whether the cases came to attention because their values were extreme.
- Toward which mean? Cases regress toward the average of whatever larger group they’re really drawn from. A strong bowler’s bad week is followed, on average, by a return toward their usual level, not toward the league’s.
- The r rule is a simplification. It describes averages under simple conditions. When two measurements have different spreads, or the relationship between them isn’t a straight line, the size of the effect differs, but the direction doesn’t.
Sources
- Francis Galton (1886). Regression towards mediocrity in hereditary stature. Journal of the Anthropological Institute of Great Britain and Ireland 15, 246–263.
- J. Martin Bland and Douglas G. Altman (1994). Statistics notes: Regression towards the mean. BMJ 308(6942), 1499.
- J. Martin Bland and Douglas G. Altman (1994). Statistics notes: Some examples of regression towards the mean. BMJ 309(6957), 780.
- Adrian G. Barnett, Jolieke C. van der Pols and Annette J. Dobson (2005). Regression to the mean: What it is and how to deal with it. International Journal of Epidemiology 34(1), 215–220.
- Veronica Morton and David J. Torgerson (2003). Effect of regression to the mean on decision making in health care. BMJ 326(7398), 1083–1084.
- Stephen M. Stigler (1996). The history of statistics in 1933. Statistical Science 11(3), 244–252.
- Harold Hotelling (1933). The Triumph of Mediocrity in Business, by Horace Secrist (book review). Journal of the American Statistical Association 28(184).
Last reviewed 2026-09-14.