Ecological fallacy
Also known as ecologic fallacy
The ecological fallacy is assuming that a relationship found in data about groups also holds for the individuals in those groups: if towns with more of one kind of resident have more of some outcome, then those residents must be the ones with the outcome. (“Ecological” here means data about groups or areas, such as states, schools or countries. It has nothing to do with the environment.)
The flaw is that group figures don’t say who within each group has which characteristic. A correlation across states is built from each state’s totals, and the same totals fit many different patterns among the people inside, including the opposite one. Groups also differ from each other in many ways at once, so a trend across them can reflect something about the places rather than anything about the individuals.
Examples
Dogs and break-ins
A neighborhood newsletter reports that towns in the county with more dog owners per household have fewer burglaries per household. “The lesson is clear: get a dog, and burglars will skip your house.”
The clear-cut case. The data describe towns, not houses. They can’t show whether the houses with dogs are the ones that weren’t burgled, and the same town figures would fit a world in which dogs make no difference to any particular house: towns with many dog owners might be rural towns with large yards, where burglaries are rarer for everyone. A dog might well deter burglars. Town totals can’t show it, because nothing in them connects a particular dog to a particular house.
The offices where remote work “boosts sales”
A company notices that its offices with a larger share of remote staff have higher average sales per employee. The CEO concludes that remote workers sell more and plans to move more of the sales staff to remote work.
In this company, the truth runs the other way for individuals. The offices with more remote staff are in big-city markets where everyone sells more, and in every office the remote staff sell less than their in-office colleagues. The group-level trend is real; it describes offices, and says nothing reliable about what happens when a salesperson goes remote. The Form section shows the arithmetic. David Freedman’s statistical account of the ecological fallacy works through a real example with the same structure: across US states, a positive correlation between two characteristics, while among individuals the association ran the other way.
Form
A version of the office example with two offices of 100 people each:
| Share remote | In-office staff: average sales | Remote staff: average sales | Office average | |
|---|---|---|---|---|
| Office A (small market) | 20% | $100k | $60k | $92k |
| Office B (big city) | 40% | $150k | $90k | $126k |
| All 200 staff | 30% | $121k | $80k | $109k |
Across offices, more remote work goes with higher sales ($92k to $126k). Within each office, remote staff sell less, and across the whole company, remote staff average $80k against $121k for everyone else. The office averages are accurate. The error is reading them as a fact about remote workers.
W. S. Robinson’s 1950 paper, the standard reference, distinguishes three correlations: the ecological correlation between group percentages or averages, the individual correlation across all the people, and the within-area correlations inside each group. He showed that the ecological correlation depends only on each group’s totals, which don’t determine the individual figures, so “there need be no correspondence” between the two.
Variants
- Ecological correlation. A correlation computed across areas, read as a correlation across people. Robinson noted that its size also depends on how the areas are drawn: the same data gave 0.773 across US states and 0.946 across the nine larger census divisions.
- Ecological regression. Estimating how each kind of person behaves from area-level data. Freedman explains that it rests on a “constancy assumption”, that the behavior of each kind of person doesn’t depend on the makeup of the area where they live, which fails in cases like the offices above.
- Confounding and aggregation bias. Freedman separates two problems: groups differ in other ways besides the one being studied (see Confounding), and even without that, results for groups can differ from those for individuals, which he calls aggregation bias.
- The reverse: the individualistic fallacy. Generalizing from relationships among individuals to relationships among groups. Subramanian and colleagues trace the term to Alker in 1969, and note that Susser warned epidemiologists about it in 1973 as the “atomistic fallacy”.
How it relates to its neighbors. The ecological fallacy resembles the Fallacy of division, which assigns a property of a whole to its parts; that entry mentions it as a closely related problem. The difference is that the ecological fallacy moves a relationship (a correlation or trend across groups) down to individuals, not a property, and the scholarly treatments used here discuss it as a problem of statistical inference rather than as a kind of division. The reversal in the office example is related to Simpson’s paradox, in which a trend in combined data reverses within subgroups.
When it isn’t an error
Group-level data are often the only data available, and Freedman notes they can offer valuable clues. Using them is sound when:
- The conclusion stays at the group level. “Offices with more remote staff have higher sales” is a true statement about offices, and some questions really are about groups, such as how a policy applied to a whole state changes the state’s rate.
- The group pattern is treated as a lead and checked with individual data. Freedman lists ecological studies that yielded important insights, including John Snow’s work on cholera, and notes that inferences from groups to individuals “may be correct, but are only weakly supported by the aggregate data.”
- The individual-level data agree. Once you know who has which characteristic, the group pattern can illustrate a relationship established another way.
The test: is the conclusion about the groups, or about the people in them, and do the data record the people?
Looks like it, but isn’t
The reservoir
A county health department notices that towns supplied by one reservoir have higher rates of a stomach illness than towns supplied by others. It tests the reservoir’s water and asks people who fell ill where their drinking water comes from.
The department starts from group data, but it doesn’t conclude anything about individuals from them. The town pattern prompts two checks that reach individuals directly: the water itself, and the sick people’s own water sources. That’s the lead checked with individual data condition.
Bike lanes and county injury rates
A state compares its counties over five years and finds that counties that built protected bike lanes saw larger drops in their rates of cyclist injuries. It offers funding to other counties to build them.
The data are about counties, and so is the conclusion: what tends to happen to a county’s injury rate when the county builds lanes. The state isn’t claiming that any particular rider who uses a lane is safer, which the county figures couldn’t show. Whether the counties that built lanes differed in other ways is a real question, but it’s a question about confounding, not the ecological fallacy. That’s the conclusion stays at the group level condition.
Why it happens
Group data are easier to get. Robinson observed that researchers used ecological correlations “simply because correlations between the properties of individuals are not available”, and that in each case “the substitution is made tacitly rather than explicitly.” Freedman adds that aggregate data are often easier to obtain than data on individuals, so ecological inferences will continue to be made.
Language also blurs the levels. “Towns with more dog owners have fewer burglaries” slides easily into “dog owners have fewer burglaries”, and an average or a rate feels like a description of a typical member. Once the sentence is about people, it’s easy to forget that the numbers never were.
How to respond
- Ask what the unit of the data is. Were people measured, or were states, schools or offices?
- Ask whether individual-level data exist. A survey or records that link each person’s characteristics can answer the question the group data only hint at.
- Ask what else differs between the groups. Places with more of one thing usually have more of many things.
- Check the pattern within groups. If it disappears or reverses inside each group, the group trend isn’t telling you about individuals.
- Keep the claim at the level of the data. “Counties with more X had more Y” is often all the data support, and often enough.
Evidence
The ecological fallacy is a flaw in inference, not an effect with a replication record, but its size has been measured wherever group and individual data describe the same people.
- Robinson (1950). Using the 1930 US census, Robinson found that the correlation between the percentage of the population that was Black and the percentage that was illiterate was 0.773 across states and 0.946 across the nine census divisions, while the corresponding individual correlation was 0.203. For foreign birth and illiteracy, the individual correlation was positive (0.118), but the ecological correlation was negative: −0.526 across states and −0.619 across divisions. He concluded that ecological correlations cannot validly be used as substitutes for individual correlations. According to Subramanian and colleagues, the term “ecological fallacy” was coined by Selvin in 1958, and Robinson’s paper had been cited over 1,150 times by 2008.
- A modern replication. Freedman repeated the exercise with 1995 survey data on income. Across the 50 states, the correlation between the share of residents born abroad and the share with high family incomes was 0.52, suggesting that people born abroad had higher incomes. Among individuals the correlation was −0.05: 35% of people born in the US had high incomes against 28% of those born abroad. An ecological regression on the state data put the figures at 29% and 85%.
- A health example. Freedman cites the finding that countries where fat is a larger part of the diet have higher death rates from breast cancer, and notes that later studies of individual-level data cast serious doubt on the link between breast cancer and fat intake.
- The critique of Robinson. Subramanian and colleagues reanalyzed Robinson’s census data with multilevel models and found substantial differences between states that individual characteristics didn’t account for. They accept Robinson’s central point but argue that analyzing only individuals is also a mistake, the individualistic fallacy, and that questions often need both levels at once.
Sources
- W. S. Robinson (1950). Ecological correlations and the behavior of individuals. American Sociological Review 15(3), 351–357 (reprinted in International Journal of Epidemiology 38(2), 337–341, 2009).
- David A. Freedman (2001). Ecological inference. International Encyclopedia of the Social & Behavioral Sciences (Elsevier), 4027–4030.
- S. V. Subramanian, Kelvyn Jones, Afamia Kaddour and Nancy Krieger (2009). Revisiting Robinson: The perils of individualistic and ecologic fallacy. International Journal of Epidemiology 38(2), 342–360.
Last reviewed 2026-09-13.