Collider bias
Also known as Berkson's paradox, Berkson's bias or Berkson's fallacy
Collider bias is a false association between two things that appears when you look only at cases selected on something both of them affect, or adjust for that thing in an analysis. The shared effect is called a collider, because in a diagram of what causes what, the arrows from both factors collide in it. If a school admits children who are strong at academics or at sports, its students can show a link between the two abilities even when there is almost none among children in general.
The flaw is that conditioning on a common effect manufactures a correlation. Among the admitted students, one who is weak at sports probably got in on academics, and one who is weak at academics probably got in on sports, so inside the school the abilities look like they trade off. It’s the mirror image of Confounding: a common cause makes two things go together, and adjusting for it removes the false link; a common effect doesn’t link them until you select or adjust on it.
Examples
The dating history
“Every really good-looking person I’ve dated turned out to be self-absorbed, and the kindest ones were never that attractive. Looks and kindness just don’t go together.”
The clear-cut case. Suppose the speaker goes on more than a date or two only with someone who clears a bar on looks and kindness combined. Someone low on kindness clears it only by being very attractive, and someone less attractive only by being very kind. So among the people the speaker has dated, the two qualities will run in opposite directions even if they’re unrelated among people in general. The observation about the speaker’s own dating history may be accurate. The conclusion about people was produced by the filter.
Berkson’s hospital
In some medical circles, surgeons are removing the gallbladder as a treatment for diabetes. A hospital checks its records and finds that cholecystitis (inflammation of the gallbladder) is significantly more common among its patients with diabetes than among patients in a comparison group, people who came in for glasses. The records seem to support a link.
This is Joseph Berkson’s own example from 1946, and the hospital tables he presented did show a significant positive correlation. He then showed that a hospital can find such a link even when none exists in the population. Each disease gives a person some chance of ending up at the hospital, and someone with two diseases has two chances, so combinations of diseases are overrepresented in a way that depends on the admission rates. In his worked example, cholecystitis, diabetes and refractive errors were unrelated among 10 million people. With admission probabilities of 0.15, 0.05 and 0.20, cholecystitis came out about twice as common in the hospital’s diabetes group as in its control group; with equal probabilities, it came out slightly less common. As Berkson wrote, the same results “would appear if the sampling were applied to randomly distributed cards instead of patients.”
Adjusting instead of selecting
A film blogger has data on every movie released in a decade. Across all of them, budget and critics’ scores are unrelated. To “control for prestige”, she adds whether a film was nominated for a major award to her analysis, and now bigger budgets go with worse reviews. She writes that big budgets make for worse films.
No film was dropped from the data, which is what makes this case easy to miss. Nominations depend on both good reviews and the money spent on campaigns. Among nominated films, a low-budget film was probably nominated for its reviews, and among films that weren’t nominated, a big-budget film probably had poor ones. Comparing films within each level of nomination builds in the same false trade-off that selecting on nomination would. Haidong Lu, Gregg Gonsalves and Daniel Westreich call this collider adjustment bias, as distinct from collider restriction bias, and argue it is better seen as overadjustment than as selection.
Form
What conditioning on a third variable C does depends on where C sits between two variables A and B:
| C is a... | Structure | Restricting to, or adjusting for, C |
|---|---|---|
| Collider | A → C ← B | Creates an association between A and B that wasn’t there |
| Common cause (confounder) | A ← C → B | Removes an association that C created |
| Mediator | A → C → B | Removes the part of A’s effect that works through C |
Two extensions matter in practice. Conditioning on something that is itself caused by a collider has the same effect; Hernán and Robins’s example is a study that ends up including only parents who aren’t grieving, which depends on whether the baby survived. And A and B needn’t affect C directly: if causes of each affect selection, the distortion carries through, a structure Gareth Griffith and colleagues note is sometimes called M-bias.
Variants
- Berkson’s bias (also Berkson’s fallacy or paradox): the hospital version. David Sackett renamed it admission rate bias in 1979. Jaapjan Snoep and colleagues note that Berkson’s original concerned two diseases, and describe an indirect form (a term they take from Flanders and colleagues) that affects studies of an exposure: the exposure is linked to a second disease that brings people to the hospital.
- Collider restriction: analyzing only the cases selected on a common effect, such as hospital patients, admitted students, people who volunteered or people who were tested.
- Collider adjustment: including a common effect as a control variable, as with the film awards.
- Selection by survival: when lasting depends on either of two qualities, the survivors show a spurious trade-off between them. See Survivorship bias, and the restaurant guide in Cum hoc ergo propter hoc.
- Splitting on something the treatment changed. In Simpson’s paradox, the checkout page that “loses in every group” is split on whether shoppers created an account, which the page itself influenced.
How it relates to selection bias. The sources draw the family lines differently. Hernán, Hernández-Díaz and Robins use “selection bias” for all biases that come from conditioning on a common effect of the exposure (or its causes) and the outcome (or its causes). Neil Pearce and Lorenzo Richiardi treat collider bias as the general phenomenon, selection bias as the kind in which the common effect is selection into a study, and Berkson’s bias as a kind of selection bias. Lu and colleagues argue that adjusting for a collider shouldn’t be called selection bias at all, and the Catalogue of Bias says only that selection bias “can sometimes be considered” a form of collider bias. Because they disagree, this site doesn’t list either as a kind of the other; see Selection bias for the wider sense of that term.
When it isn’t an error
Restricting or adjusting is fine, or the bias negligible, when:
- The question is about the selected group itself. An association among hospital patients is a real fact about the hospital’s patients, useful for planning its care.
- Only one of the two factors affects selection. Berkson noted that if a characteristic has no influence on who comes to the hospital, as with eye color, comparing it across disease groups there is not distorted in this way.
- The variable isn’t a common effect. Adjusting for something that comes before both factors and influences them, such as age, is adjusting for a possible confounder, not a collider.
- The distortion is too small to matter. Snoep and colleagues conclude that when studies use newly diagnosed cases, Berkson’s bias becomes highly unlikely for exposure-disease questions, barring unusual circumstances, and Pearce and Richiardi stress that showing a bias could occur isn’t the same as showing it is large.
The test: did both things being compared influence which cases got into the data, or the category you adjusted for?
Looks like it, but isn’t
Adjusting for age
A study of whether daily coffee affects sleep quality adjusts for age. A reviewer objects that adjusting for extra variables can create collider bias and asks for the unadjusted results only.
Adjustment can create collider bias, which is why the objection sounds careful. But age isn’t caused by coffee or by sleep, so it can’t be their common effect; it plausibly affects both how much coffee people drink and how they sleep, which makes it a candidate confounder. Pearce and Richiardi warn against this kind of “collider anxiety”, the fear of adjusting for a real confounder in case it is a collider. That’s the isn’t a common effect condition.
Planning care on the ward
A hospital’s planners notice that among inpatients, diabetes and gallbladder disease occur together more often than population rates would predict, so they arrange for the diabetes team to review patients before gallbladder surgery.
This is exactly the pattern Berkson warned about, but the planners aren’t concluding that one disease causes the other or that the link exists outside the hospital. They’re describing the patients they actually treat, and for that the hospital data are the right data. That’s the question about the selected group condition.
Why it happens
Selection is usually invisible in the finished data set. The hospital’s records look like “the data”, not like a filtered sample, and nothing in them shows the people who never came in. And the patterns selection creates are often easy to explain: shabby restaurants spending on food, attractive people becoming complacent. A plausible story makes the correlation feel discovered rather than manufactured.
Conditioning also feels like caution. Restricting to comparable cases and “controlling for” more variables are habits of careful analysis, and they are the right move for confounders, so it’s natural to apply them to everything. Berkson made the underlying point in 1946: the spurious correlations he described require no biological connection at all, only “the ordinary compounding of independent probabilities.”
How to respond
- Ask what got each case into the data. If both factors being compared could influence it, expect a distorted association.
- Draw the arrows. For any variable a study restricts or adjusts for, ask whether it is a cause of the two factors or an effect of them. Causal diagrams exist for exactly this.
- Look for a sample not selected on the common effect. The Catalogue of Bias describes an analysis of the same two diseases in hospital patients and in the general population that gave very different answers (see Evidence).
- Ask how big the bias could plausibly be. A collider that depends only weakly on the factors produces little distortion, and Pearce and Richiardi warn that fear of collider bias can paralyze analysis.
- At the design stage, keep the exposure and outcome from driving inclusion. The Catalogue of Bias recommends inclusion criteria chosen so that neither influences who is enrolled or stays in a study.
Evidence
Collider bias is a structural feature of how data are selected or analyzed, not an effect with a replication record. Its size has been checked where data on the unselected population exist.
- Berkson’s argument was mathematical. His 1946 paper used a constructed population. According to Snoep and colleagues, R. S. Roberts, David Sackett and colleagues set out to test it in 1978 using household surveys, and found a distorted association, larger or smaller, for most of 28 pairs of diseases among people who had been hospitalized in the previous six months. The Catalogue of Bias reports one pair: locomotor and respiratory disease had an odds ratio of 4.06 among 257 hospitalized people and 1.06 among 2,783 people in the general population.
- It may rarely have mattered in its classic form. Snoep and colleagues conclude that the original fallacy is a probabilistic necessity in hospital studies comparing diseases people already have, but that the usual design choices of hospital-based case-control studies seem to preclude a large role for it in epidemiology. Pearce and Richiardi call it, in practice, “much ado about nothing”, while showing numerically that its direction depends on how controls are chosen.
- Non-random samples. Griffith and colleagues (2020) argued that studies of COVID-19 risk factors based on hospitalized patients, people who had been tested, or volunteers were open to collider bias. In the UK Biobank cohort, they found that 811 of 2,556 characteristics examined (32%) were associated with whether a participant had been tested, and they recommend addressing the problem through sampling strategy at the design stage.
Sources
- Joseph Berkson (1946). Limitations of the application of fourfold table analysis to hospital data. Biometrics Bulletin 2(3), 47–53 (reprinted in International Journal of Epidemiology 43(2), 511–515, 2014).
- Miguel A. Hernán and James M. Robins (2020). Causal Inference: What If (chapters 6 and 8). Chapman & Hall/CRC (online edition revised 2026).
- Miguel A. Hernán, Sonia Hernández-Díaz and James M. Robins (2004). A structural approach to selection bias. Epidemiology 15(5), 615–625.
- Jaapjan D. Snoep, Alfredo Morabia, Sonia Hernández-Díaz, Miguel A. Hernán and Jan P. Vandenbroucke (2014). Commentary: A structural approach to Berkson's fallacy and a guide to a history of opinions about it. International Journal of Epidemiology 43(2), 515–521.
- Neil Pearce and Lorenzo Richiardi (2014). Commentary: Three worlds collide: Berkson's bias, selection bias and collider bias. International Journal of Epidemiology 43(2), 521–524.
- Haidong Lu, Gregg S. Gonsalves and Daniel Westreich (2023). Selection bias requires selection: The case of collider stratification bias. American Journal of Epidemiology 193(3), 407–409.
- H. Lee, Jeffrey K. Aronson and David Nunan (2019). Collider bias. Catalogue of Bias.
- E. A. Spencer, Jeffrey K. Aronson, David Nunan and C. Heneghan (2018). Admission rate bias (Berkson's bias). Catalogue of Bias.
- Gareth J. Griffith, Tim T. Morris, Matthew J. Tudball and others (2020). Collider bias undermines our understanding of COVID-19 disease risk and severity. Nature Communications 11, 5749.
Last reviewed 2026-09-13.