Confounding
Also known as confounding bias
Confounding happens when something influences both an exposure (a treatment, habit or condition being studied) and an outcome. That third factor, a confounder, makes the exposure and the outcome go together, or fail to, for reasons that have nothing to do with the exposure’s own effect. Hernán and Robins’s illustration: whether one pedestrian’s looking up makes others look up is hard to judge when a thunderous noise overhead makes everyone look up.
The flaw is that the groups being compared differed before the exposure did anything. People who got the exposure were, on average, different kinds of people in ways that also affect the outcome, so comparing them with everyone else measures those differences along with any effect. The Catalogue of Bias puts the consequence simply: confounding can suggest an association where none exists, or mask a true one. It is a flaw in how a comparison is set up, so it can mislead no matter how large the study is.
Examples
Shoe size and reading
A school district analyzes test results for all its elementary students and finds a strong correlation: children with bigger feet score higher in reading. A board member wonders aloud whether the district should look into it.
The clear-cut case. Age drives both: older children have bigger feet, and older children read better. Within a single grade the correlation would largely disappear. Nobody would be fooled here, which is what makes it useful; the same structure is at work in the harder cases below, where the third factor is less obvious.
The fungicide that “causes” brown patches
A lawn-care company’s records show that lawns treated with its fungicide were more likely to have brown patches by August than untreated lawns. A gardening blog concludes that the fungicide damages grass.
The company treats lawns whose owners call in because they’re already seeing signs of fungal disease. Early disease leads to treatment and also, independently, to brown patches later in the summer, so treated lawns were headed for worse outcomes before any fungicide was applied. This is confounding by indication: the reason for giving a treatment is itself a risk factor for the outcome. Hernán and Robins’s medical version is aspirin, which is prescribed more often to people with heart disease, who are also at higher risk of stroke. It can make a treatment that helps look useless or harmful.
The routine that “doesn’t work”
A running club surveys its members. Those who do a weekly strength-training routine report injuries just as often as those who don’t, so the club’s newsletter says the routine doesn’t prevent injuries. A member points out that the strength trainers run more miles; the newsletter reruns the comparison among runners with similar weekly mileage and still finds no difference.
Here confounding may be hiding an effect instead of creating one. Runners who log high mileage are more likely both to take up strength training and to get injured, which pushes the injury rates of the two groups together even if the routine helps. Adjusting for mileage deals with that. But many runners start a routine after an injury, on a physiotherapist’s advice, and injury history also predicts future injury. The survey didn’t ask about it, and no adjustment can account for a factor that wasn’t measured. The data can’t say whether the routine helps, harms or does nothing.
Form
Hernán and Robins define confounding by its causal structure: a common cause of the exposure and the outcome. What adjusting for a third variable does depends on where that variable sits:
| Third variable C | Structure | Adjusting for C |
|---|---|---|
| Confounder | Exposure ← C → Outcome | Removes the misleading association it creates (if C is measured well) |
| Mediator | Exposure → C → Outcome | Removes part of the real effect, the part that works through C |
| Collider | Exposure → C ← Outcome (or ← causes of each) | Creates an association that wasn’t there |
So “controlling for more variables” isn’t automatically safer. A variable that the exposure itself affects is usually a mediator or a collider, not a confounder, and adjusting for it is an error of its own (see Collider bias and the look-alikes below). Hernán and Robins also note that economists often call confounding “omitted variable bias”.
Variants
- Confounding by indication. The reason a treatment was given also affects the outcome, as with the fungicide. Hernán and Robins note a related term, channeling, for patient risk factors that lead doctors to choose one drug within a class.
- Healthy worker bias. Hernán and Robins’s example: studying whether working as a firefighter affects the risk of death is confounded if physical fitness makes people both more likely to be active firefighters and less likely to die. (The same name is also used for a related selection problem.)
- Reverse causation in disguise. An early, undiagnosed stage of a disease can cause both an exposure and the outcome, as when subclinical illness leads people to exercise less and later becomes clinical disease. Hernán and Robins note this is often called reverse causation when the underlying factor is unknown.
- Residual or unmeasured confounding. A study adjusts for the confounders it measured, but others remain. The Catalogue of Bias stresses that there will always be the possibility of unknown confounders.
- Exaggerating or masking. Depending on how the confounder relates to the exposure and the outcome, it can push an estimate up or down, as in the running club, where it may hide a benefit.
- Reversal. When a confounder is strong enough, the association can point the opposite way from the effect. Julious and Mullee call Simpson’s paradox “an extreme example” of confounding, “in which this third factor reverses the effect first observed”. (Simpson’s paradox can also come from splitting data on the wrong kind of variable, which that entry discusses.)
How this fits with its neighbors. Confounding is the research version of one alternative explanation that a correlation leaves open; the correlation vs. causation page lists the others. Concluding a cause from a correlation without considering them is cum hoc ergo propter hoc, or, from a sequence in time, post hoc ergo propter hoc. Hernán and Robins separate confounding (bias from common causes) from selection bias (bias from conditioning on a common effect, the collider structure), while noting that not every discipline draws the distinction this way.
When it isn’t an error
Confounding is a threat to interpreting an association as an effect. It doesn’t apply, or can be set aside, when:
- The exposure was randomly assigned. A coin flip can’t be influenced by anything that affects the outcome, so in a randomized experiment no common cause is expected. Chance imbalances can still occur, and the Catalogue of Bias notes that poorly designed randomized studies can be confounded.
- The factors that decided exposure were measured and adjusted for. If a set of variables that come before the exposure accounts for every common cause, adjusting for them removes the confounding. Hernán and Robins stress that this is an assumption that can’t be checked from the data and rests on prior knowledge of what causes what.
- The question is prediction, not effect. If you only want to predict who will have brown patches, “was treated with fungicide” is a legitimately useful predictor. Confounding matters when the association is read as what would happen if you changed the exposure.
- The effect is too large for plausible confounders. The Catalogue of Bias notes that a very large effect can outweigh the combined effects of plausible confounders, citing general anesthesia, whose effects are unlikely to be explained by confounding. Hernán and Robins add that a large bias requires a confounder strongly associated with both exposure and outcome.
- The proposed “confounder” isn’t one. A factor the exposure itself affects (a mediator, or a collider) doesn’t confound the total effect, and adjusting for it would introduce bias.
The test: does this third factor influence both who got the exposure and the outcome, without itself being a result of the exposure?
Looks like it, but isn’t
Not adjusting for plant height
A seed company runs a trial in which tomato plants are randomly assigned to a new fertilizer or to none. Fertilized plants yield 20% more fruit. A reviewer objects that the analysis didn’t control for plant height: “Taller plants produce more fruit, and the fertilized plants were taller. Once you adjust for height, the advantage mostly disappears.”
The reviewer has found a factor associated with both the fertilizer and the yield, which is how confounders are often spotted, so the objection sounds like good methodology. But height isn’t a common cause of getting fertilizer and yielding more; the fertilizer caused the extra height, which is one of the ways it raises yield. Height is a mediator. Adjusting for it asks how much fertilizer helps plants of the same height, which removes part of the very effect the trial set out to measure. Hernán and Robins call this overadjustment for mediators when the total effect is the question. Declining to adjust is correct, and random assignment already rules out confounding: that’s the isn’t one and randomly assigned conditions.
Not adjusting for doctor visits
A company randomly assigns 400 office workers to standing desks or ordinary desks and compares back pain scores after six months. A critic says the comparison should be restricted to, or adjusted for, whether each worker saw a doctor during the study, “since people who see doctors are different”.
People who see doctors are different, which is why this sounds like a confounding concern. But seeing a doctor happened after assignment and is influenced both by the outcome (back pain sends people to doctors) and possibly by the desk (a new standing desk might prompt a checkup for sore feet). A variable affected by both is a collider, and restricting or adjusting for it can create a difference between the groups that the desks didn’t cause. Hernán and Robins show that conditioning on a common effect of treatment and outcome produces this kind of bias even in a randomized experiment. The unadjusted comparison is the right one. See Collider bias for how this works.
Why it happens
In observational data, nobody flips a coin to decide who is exposed. People choose their habits, doctors choose treatments for the patients who seem to need them, and circumstances such as age, money or health shape both what people are exposed to and what happens to them. Hernán and Robins describe the result: the effects of those factors “become entangled with the effect of treatment.”
It is easy to miss for two reasons. The confounder is often not in the data at all, so nothing in the analysis draws attention to it. And the natural fix, adjusting for everything available, can make things worse. Hernán and Robins show that the traditional way of picking confounders, by checking which variables are statistically associated with both exposure and outcome, can lead researchers to adjust for variables that aren’t confounders and so introduce bias; deciding what to adjust for requires a view of what causes what.
How to respond
- Ask how the exposure was decided. Who or what determined who got it? Anything that shaped that decision and also affects the outcome is a candidate confounder.
- Name a specific confounder and its likely direction. Hernán and Robins note that the direction of the bias can often be reasoned out: if smokers are less likely to receive a treatment and more likely to die, failing to adjust for smoking makes the treatment look better than it is.
- Check when each adjustment variable was measured. Factors fixed before the exposure are candidates for adjustment; factors the exposure could have changed usually aren’t.
- Ask whether the size of the effect could plausibly be explained. A small association is easy for a modest confounder to produce; a very large one needs a confounder that is strongly tied to both exposure and outcome. Sensitivity analyses, which recompute the result under different assumptions about unmeasured confounders, formalize this.
- Look for randomized or natural-experiment evidence on the same question.
- Don’t use “could be confounded” as a blanket dismissal. Every observational study could be. The objection has force when it names a plausible factor strong enough to matter; the correlation vs. causation page discusses this misuse.
Evidence
Confounding is a feature of study design, not an effect with a replication record, but its consequences have been documented where observational findings were later tested.
- Observational results overturned by trials. The Catalogue of Bias describes early findings of a beneficial effect of hormone replacement therapy on cardiovascular disease that were reversed once studies adjusting for socioeconomic status or education were taken into account; the reduced risk appeared among the studies that hadn’t adjusted for them. It also describes retrospective, non-randomized studies of the heart drug digoxin that found higher death rates even after adjustment, while a prospective randomized study found no increase in mortality; people taking digoxin were likely sicker to begin with.
- The effect of one missing adjustment. In a systematic review of observational studies of statins and Parkinson’s disease cited by the Catalogue of Bias, six studies that didn’t adjust for cholesterol (which is inversely related to the risk of Parkinson’s) together showed a protective effect (relative risk 0.75). In the four that did, the estimate shifted to a non-significant 1.04.
- Strength as a defense. Austin Bradford Hill contrasted smoking’s association with lung cancer, where death rates in smokers were nine to ten times those in nonsmokers, with its association with coronary thrombosis, no more than twice. For the smaller association, he wrote, it is “much easier ... to think of some features of life that may go hand-in-hand with smoking” that might be the real cause; for the larger one, a confounder would have to be so closely linked with smoking that it “should be easily detectable”.
Sources
- Miguel A. Hernán and James M. Robins (2020). Causal Inference: What If (chapters 3, 7, 8, 18 and 22). Chapman & Hall/CRC (online edition revised 2026).
- Jeffrey K. Aronson, Clare Bankhead and David Nunan (2018). Confounding. Catalogue of Bias.
- Steven A. Julious and Mark A. Mullee (1994). Confounding and Simpson's paradox. BMJ 309(6967), 1480–1481.
- Austin Bradford Hill (1965). The environment and disease: Association or causation?. Proceedings of the Royal Society of Medicine 58(5), 295–300.
Last reviewed 2026-09-13.