Skip to content
Check My Logic
Check My Logic

Principle

Correlation vs. causation

Also known as correlation does not imply causation or association is not causation

Correlation (or association) means two things tend to go together: when one is higher, more common or present, the other tends to be too, or tends to be lower. Causation means one of them makes a difference to the other: if you changed the first, the second would change. The principle is that the first doesn’t establish the second. Two things can rise and fall together for reasons that have nothing to do with one acting on the other.

This is a methodological principle, a rule about what a kind of evidence can and can’t show. It isn’t a logical law, and it isn’t a single effect measured in an experiment. It also doesn’t say correlations are worthless: as Hill observed of occupational medicine, more often than not the first sign of a hazard is an association between an illness and something in the environment.

Example

A city report notes that neighborhoods with more coffee shops have higher rents. A council member proposes subsidizing new coffee shops in low-rent neighborhoods to raise property values.

The correlation may be perfectly real, but it fits several different stories. Coffee shops might raise rents by making a neighborhood more attractive. High rents might attract coffee shops, because they open where customers have money to spend (the effect running the other way). Something else, such as how close a neighborhood is to downtown, might drive both. Only the first story supports the subsidy, and the correlation alone can’t say which story is true, or how much of each.

What “causes” means here

The philosopher David Lewis stated the guiding idea of counterfactual theories of causation this way, as quoted in the Stanford Encyclopedia of Philosophy: “We think of a cause as something that makes a difference, and the difference it makes must be a difference from what would have happened without it.” The epidemiologists Miguel Hernán and James Robins make the same contrast in statistical terms. An association compares two different groups as they actually are: people who took a treatment against people who didn’t. A causal effect compares the same group under two conditions: everyone treated against everyone untreated. For any one person, only one of those conditions can actually be observed. The two answers can differ whenever the groups differed to begin with.

The statistician Austin Bradford Hill, writing about occupational disease, put the practical question plainly: whether “the frequency of the undesirable event B will be influenced by a change in the environmental feature A.”

Why things can go together without one causing the other

When A and B are correlated, the possibilities include the following (the first four follow a standard list that Martijn Demollin takes from the logic textbook author Trudy Govier):

  • A causes B. The correlation is what it looks like.
  • B causes A (reverse causation). Hill’s example: does a particular diet lead to a disease, or do the early stages of the disease lead to that diet? He called this the question of “which is the cart and which the horse?”
  • Something else causes both (a common cause, or confounder). Hernán and Robins’s example: a researcher wants to know whether one pedestrian looking up at the sky makes others look up, but a thunderous noise overhead makes everyone look up, so it’s unclear what the second pedestrian was reacting to. See Confounding.
  • Coincidence. With enough variables compared, some will line up by chance. Looking through many comparisons and reporting the ones that did is the problem behind P-hacking and the Texas sharpshooter fallacy.
  • The way cases were selected. If the data include only cases selected on something that both A and B influence, A and B can become associated even when neither affects the other. Hernán and Robins call this conditioning on a common effect; see Collider bias and Selection bias.

These can also combine: a real effect plus a confounder that exaggerates it, or a confounder that hides a real effect. The Catalogue of Bias notes that confounding can suggest an association where none exists or mask a true one. So the absence of a correlation doesn’t prove the absence of a cause either, and a correlation can even reverse direction when the data are split, as in Simpson’s paradox.

Where the errors live on this site

The principle is the common thread of several entries, each about a different way of getting it wrong:

  • Post hoc ergo propter hoc: B followed A, so A caused B. The error of reasoning from sequence in time.
  • Cum hoc ergo propter hoc: A and B occur together, so one causes the other. The error of reasoning from co-occurrence alone, including getting the direction wrong.
  • Confounding: the research bias in which a third factor that affects both the exposure and the outcome distorts a study’s estimate of an effect, and what adjusting for it can and can’t do.
  • The regression fallacy: crediting whatever happened in between for a change that was really a return toward average.
  • Illusory correlation: perceiving a correlation that isn’t in the data at all, a step before any of the above.

What does help

No single test turns a correlation into a proven cause. What the methods below have in common is that each rules out some of the alternative explanations listed above.

Experiments

In a randomized experiment, a coin flip or a random number generator decides who gets the treatment. Hernán and Robins explain why that matters: nothing that affects the outcome can also have decided who was treated, because a coin flip has no such influence, so the treated and untreated groups are expected to be comparable; in their words, “in ideal randomized experiments, association is causation.” In the Stanford Encyclopedia’s account of causal models, Christopher Hitchcock describes randomized trials as aiming at an intervention in James Woodward’s sense: a causal process that operates independently of the other variables in the model. In a trial of a blood pressure drug, factors like education and health insurance that normally influence who takes the drug no longer do. Randomization doesn’t remove chance: the Catalogue of Bias notes that groups can still end up imbalanced by chance, and that confounding can occur in randomized studies, especially poorly designed ones.

Natural experiments

When an experiment isn’t possible, sometimes the world supplies something close to one. Peter Craig and colleagues, following UK Medical Research Council guidance, define a natural experiment broadly as any event not under a researcher’s control that divides a population into exposed and unexposed groups, such as a new law or a change in who is eligible for a program. They trace the approach back to John Snow’s studies of cholera in London; Hill describes one of Snow’s comparisons from 1854, in which houses supplied with grossly polluted water by one company had 14 times the cholera death rate of houses supplied with sewage-free water by its rival. Craig and colleagues stress the main weakness, selective exposure: the people exposed may differ from the unexposed in ways that affect the outcome, so understanding how exposure came about is central to the design.

Adjusting for known alternatives

Observational studies can compare like with like by adjusting for factors that might drive both exposure and outcome. Hernán and Robins emphasize two limits. Adjustment needs prior causal knowledge of which factors to adjust for, since adjusting for the wrong kind of variable can create bias rather than remove it (see Confounding). And it can’t account for factors nobody measured. The Catalogue of Bias notes that it is not uncommon for observational findings to be overturned by later randomized trials for this reason.

Hill’s viewpoints

In a 1965 lecture, Hill listed nine aspects of an association to consider “before deciding that the most likely interpretation of it is causation”. They are often called the Bradford Hill criteria, but Hill called them “viewpoints” and explicitly declined to make them rules: “None of my nine viewpoints can bring indisputable evidence for or against the cause-and-effect hypothesis and none can be required as a sine qua non.” In brief:

Viewpoint The question it asks
Strength How large is the association? A large one is harder to explain away by some other factor.
Consistency Has it been seen repeatedly, by different people, in different places, circumstances and times?
Specificity Is it limited to particular exposures and particular outcomes? Hill warned against over-emphasizing this one.
Temporality Did the supposed cause come first?
Biological gradient Does more exposure go with more of the outcome (a dose–response relationship)?
Plausibility Does it make sense given current knowledge? Hill: “this is a feature I am convinced we cannot demand.”
Coherence Does a causal reading seriously conflict with what is known of the disease’s natural history and biology?
Experiment When the exposure is reduced or removed, does the outcome change? Hill thought this might give “the strongest support”.
Analogy Have similar causes been shown to have similar effects?

Hill’s dose–response example was that lung cancer death rates rose with the number of cigarettes smoked daily, which “adds a very great deal” to the simple comparison of smokers with nonsmokers. Evidence can also cut the other way: the Catalogue of Bias describes a study in which an apparent link between a class of antidepressants taken in pregnancy and certain birth defects showed no dose–response relationship, and similar risks appeared in women who had stopped taking the drugs during pregnancy. The Catalogue takes these as evidence of an unidentified confounder, not a true association.

Hill’s own summary question was: “is there any other way of explaining the set of facts before us, is there any other answer equally, or more, likely than cause and effect?”

Limits

A correlation is still evidence. If A causes B, A and B will usually be correlated, so finding the correlation is what the causal hypothesis predicts, and failing to find it would count against it. Hernán and Robins point out that associations produced by a common cause are real associations; what they can’t be is interpreted as effects. Hill’s nickel refiners are a case where statistics carried the argument almost alone: among about 1,000 workers and pensioners of a South Wales nickel refinery, sixteen died of lung cancer between 1929 and 1938 where about one death was expected, and eleven died of nasal sinus cancer where a fraction of one was expected. No causal agent had been identified, yet Hill asked whether anyone would “hesitate to accept it as proof of a grave industrial hazard”.

The slogan can be misused to dismiss good evidence. “Correlation isn’t causation” is a reason to ask what else could explain the correlation, not a refutation. Hill rejected what he called “the vague contention of the armchair critic ‘you can’t prove it, there may be such a feature’”: when an association is very strong, a hidden factor able to explain it would have to be so closely tied to the exposure that it “should be easily detectable”. Hernán and Robins make a related point: a large bias from confounding requires a confounder strongly associated with both the exposure and the outcome, so a critic’s alternative can be weighed for whether it could produce an effect of the size observed. An objection that names no alternative, or names one too weak to matter, isn’t a reason to set the evidence aside. Refusing any causal conclusion short of a perfect experiment has the shape of the nirvana fallacy. As Hill put it, “All scientific work is incomplete”, and that “does not confer upon us a freedom to ignore the knowledge we already have”.

Spotting the gap doesn’t settle the question. Pointing out that someone has argued from correlation alone shows their argument is incomplete. It doesn’t show the causal claim is false; the cause may be real and simply not yet shown.

Sources

  1. Austin Bradford Hill (1965). The environment and disease: Association or causation?. Proceedings of the Royal Society of Medicine 58(5), 295–300.
  2. Miguel A. Hernán and James M. Robins (2020). Causal Inference: What If (chapters 1, 2, 3, 7 and 8). Chapman & Hall/CRC (online edition revised 2026).
  3. Christopher Hitchcock (2018). Causal models. Stanford Encyclopedia of Philosophy.
  4. Peter Menzies and Helen Beebee (2024). Counterfactual theories of causation. Stanford Encyclopedia of Philosophy (substantive revision).
  5. Peter Craig, Srinivasa Vittal Katikireddi, Alastair Leyland and Frank Popham (2017). Natural experiments: An overview of methods, approaches, and contributions to public health intervention research. Annual Review of Public Health 38, 39–56.
  6. Martijn H. Demollin (2021). The argument from correlation to cause in science communication. Argument: Biannual Philosophical Journal 10(2), 315–331.
  7. Jeffrey K. Aronson, Clare Bankhead and David Nunan (2018). Confounding. Catalogue of Bias.

Last reviewed 2026-09-13.