Texas sharpshooter fallacy
Also known as Texas sharpshooter
The Texas sharpshooter fallacy takes its name from a proverbial marksman who fires a few shots at the side of a barn, then paints a target around the tightest group of bullet holes and points to it as proof of his aim. In reasoning, it’s noticing a cluster or pattern in data, drawing the boundaries around it after the fact, and then treating the pattern as if it had been predicted.
The flaw is that a target drawn around the holes can’t be missed. In any large set of data, some clusters will turn up somewhere by chance. How unlikely this particular cluster looks is the wrong question, because nobody named it in advance; the relevant question is how likely it was that some cluster would appear, somewhere, among all the places one could have looked. A pattern found this way is a lead worth checking, not a confirmation.
Examples
The third-floor printer
Four people who sit on the third floor of an office have had frequent headaches this year. Someone sketches the floor plan and notices that all four sit within 20 feet of the new printer. “That can’t be a coincidence. It’s the printer.”
The boundaries were drawn around the cases: this floor, this year, this radius, chosen because they contain the four people. Across all the floors, desks, possible causes and time periods someone could have examined, a group like this would likely turn up somewhere by chance. Epidemiologists call one version of this boundary shrinkage: drawing the edges of a suspected cluster tightly around the known cases, which inflates the apparent rate. None of this shows the printer is harmless. It shows the pattern can’t yet tell anyone whether it is.
The subgroup that came out significant
A company surveys 2,000 customers about a new logo, asking 20 questions. An analyst compares answers across 10 demographic groups and finds that left-handed customers over 50 like the logo much more than everyone else, a statistically significant difference. The report leads with “Logo has strong appeal among older left-handed customers.”
With 20 questions and 10 groups there are 200 comparisons, and at the usual 5% threshold, about 10 of them would look “significant” by chance even if the logo appealed to everyone equally. The report picks the one that did and presents it as a finding. Researchers call related practices p-hacking (trying analyses until one comes out significant) and HARKing (hypothesizing after the results are known: presenting a hypothesis formed from the results as if it had been set out beforehand).
The batches that “didn’t count”
A home baker tests whether resting bread dough overnight makes better bread. Of six rested loaves, four taste better than the unrested ones and two don’t. Afterward she decides that the oven must have been running cool on the day she baked those two, and leaves them out: “Four out of four. Resting works.”
This is the target moved in the other direction. Instead of drawing the bullseye around the hits, she redraws what counts as a valid shot so the misses don’t count. The reason for excluding the two loaves was found only after she knew they had disappointed. Kleber Neves and Olavo Amaral call this the reverse Texas sharpshooter: throwing out experiments that “didn’t work” on grounds that seem critical only after the results are in.
Form
The difference between a test and a Texas sharpshooter is the order of the steps:
| A real test | Texas sharpshooter | |
|---|---|---|
| First | Say what you expect to find, where, and what would count | Collect the data |
| Then | Collect the data | Find where the data cluster |
| Then | See whether the data hit the target | Draw the target around the cluster |
| What it shows | Whether the prediction held | Nothing yet: the target was fitted to the data |
A post hoc (after-the-fact) pattern isn’t worthless. It becomes evidence once it’s tested against data that played no part in drawing it.
Variants
- Drawing boundaries around a cluster: in space, time or group membership. The classic setting is investigating a reported disease cluster. The epidemiologist Kenneth Rothman, writing in 1990 about investigations of reported disease clusters, invoked, as Goodman and colleagues quote him, the proverbial sharpshooter “who first fires his bullet and only then draws the target”.
- Hunting through many comparisons and reporting the ones that came out significant. Torsten Biemann uses the name for researchers who keep significant predictors in a statistical model and drop the rest, and shows by simulation how this distorts later summaries of the research.
- Presenting a hypothesis formed after the results as if it came first (HARKing). Dan Lewer and colleagues describe the sharpshooter as an analogy for exactly this.
- Shifting the criteria for a match once the answer is known. William Thompson documents forensic DNA analysts shifting what counts as a “match” after learning a suspect’s profile, which makes matches look more telling than they are.
- The reverse sharpshooter: excluding results after seeing them, as in the bread example.
Not the same error: the clustering illusion. The clustering illusion is the tendency to see streaks and clumps in random data as meaningful. The Texas sharpshooter fallacy is the reasoning step that can follow: drawing the boundary around the clump and treating it as confirmation. On this site’s usage, one is a perception and the other an argument; they often go together.
When it isn’t an error
Noticing patterns after the fact is how many investigations start. It holds up when:
- The target was set before the data were seen. A prediction written down in advance, such as a preregistered study, can’t be moved to fit the results. For forensic DNA work, Thompson calls for procedures such as “sequential unmasking” that fix the target before the shots are taken.
- The pattern is treated as a hypothesis and tested on new data. Olsen, Martuzzi and Elliott note that cluster analyses, like single case reports, can generate new knowledge, even though they rarely identify causes on their own.
- The number of places looked is taken into account. Statistical corrections for multiple comparisons, or testing for clustering across a whole map rather than around one suspicious spot, build the “somewhere” back into the calculation.
- An unusual, specific link was there independently. Goodman and colleagues describe the most informative cancer cluster studies as those whose cases shared an occupation or an unusual risk factor, such as a cluster of a rare cancer among people who worked with asbestos. The shared exposure isn’t a boundary drawn around the cases; it’s a distinct explanation that can be checked.
- Exploration is labeled as exploration. Reporting “we noticed this and it needs testing” claims only what the data can support.
The test: was the target drawn before or after the shots, and has it been checked against shots fired since?
Looks like it, but isn’t
Three sick dogs, one brand of treats
A veterinarian sees three dogs in one week with the same rare set of kidney symptoms, a condition she sees perhaps once a year. Asking the owners, she learns all three recently switched to the same new brand of treats. She reports the brand to the regulator and asks for the treats to be tested.
The cluster was noticed after the fact, but two things set it apart. The condition is rare enough that three cases in a week is far outside anything chance would normally produce in her practice, and the explanation isn’t a boundary drawn around the dogs. It’s a specific, shared exposure with a plausible biological route, which the vet checked for rather than inferred from where the cases fell. Most importantly, she treats it as a lead to be tested, not a conclusion. That’s an unusual, specific link plus tested on new data.
A pattern checked against next month
A café owner looks through last year’s sales and notices that iced drinks sell noticeably better on days above 80°F. Before trusting it, he writes down his expectation for the coming month’s forecasted hot days, then compares it with what actually sells.
He found the pattern by looking at old data, like the sharpshooter. But he then fixed the target in advance and checked it against data that played no part in finding it. If the pattern was a fluke of last year, next month is unlikely to show it. That’s the tested on new data condition.
Why it happens
People are quick to see patterns and slow to ask how many places they looked. After the fact, the boundaries that were chosen tend to feel natural: “the third floor” is a real place and “customers over 50” is a real group, so it’s easy to forget they were picked because the cases fell inside them. Once a pattern suggests an explanation, confirmation bias makes evidence for it stand out and evidence against it look like noise to be explained away, which is how the reverse sharpshooter happens even to careful researchers.
The probabilities are also genuinely counterintuitive. The chance that a particular office has four cases is small; the chance that some office somewhere does is large. Michael Goodman and colleagues cite an earlier comparison between calculating the odds for a group already known to be unusual and choosing lottery numbers after the draw.
A pattern chosen because it was extreme will usually look less extreme when new data come in, since part of what made it stand out was chance. Mistaking that fade for the effect of something done in the meantime is the Regression fallacy.
How to respond
- Ask when the boundary was drawn. Were “the third floor”, “this year” or “left-handed customers over 50” chosen before looking, or because that’s where the cases were?
- Ask how many other places could have been searched. How many floors, groups, time periods, questions or outcomes were available? The more there were, the less a single cluster means.
- Ask for new data. The fair test of a pattern found after the fact is whether it shows up again in data that weren’t used to find it.
- Don’t dismiss the pattern either. Pointing out the fallacy shows the evidence is weak, not that the suspected cause is harmless. The printer might really be the problem.
- In your own work, write the target down first: what you expect, what would count, and what you’d exclude and why, before seeing the results.
Evidence
The Texas sharpshooter fallacy is a flaw in reasoning and method, not an effect with a replication record, but its costs have been studied in the fields where it is most at home.
- Disease clusters. At a 1989 conference on disease clusters, Rothman (1990) argued that investigations of individual reported clusters rarely advance understanding, listing among the reasons that the population of interest is selected by after-the-fact reasoning. Goodman and colleagues (2012) reviewed 428 US cancer cluster investigations since 1990. An increase in incidence was confirmed for 72 of the 567 cancer categories examined (13%), three were linked to a hypothesized exposure with varying certainty, and only one investigation revealed a clear cause. They attribute this to several methodological problems, of which after-the-fact boundaries are only one.
- Forensic science. Thompson (2009) used casework, informal and naturalistic experiments and analysts’ own testimony to show how shifting match criteria after seeing a suspect’s profile distorts the statistics used to describe DNA matches.
- Research analysis. Biemann (2013) showed in simulations that keeping only significant predictors in published regressions can inflate effect sizes when studies are combined in meta-analyses, especially when samples are small (under 100). Lewer and colleagues (2025) argue that a common medical study design, testing many “factors associated with” an outcome and interpreting whichever are significant, combines this fallacy with several related practices.
Sources
- Kenneth J. Rothman (1990). A sobering start for the cluster busters' conference. American Journal of Epidemiology 132(suppl. 1), S6–S13.
- Sjurdur F. Olsen, Marco Martuzzi and Paul Elliott (1996). Cluster analysis and disease mapping: Why, when, and how? A step by step guide. BMJ 313(7061), 863–866.
- Michael Goodman, Joshua S. Naiman, Dina Goodman and Judy S. LaKind (2012). Cancer clusters in the USA: What do the last twenty years of state and federal investigations tell us?. Critical Reviews in Toxicology 42(6), 474–490.
- William C. Thompson (2009). Painting the target around the matching profile: The Texas sharpshooter fallacy in forensic DNA interpretation. Law, Probability and Risk 8(3), 257–276.
- Torsten Biemann (2013). What if we were Texas sharpshooters? Predictor reporting bias in regression analysis. Organizational Research Methods 16(3), 335–363.
- Norbert L. Kerr (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review 2(3), 196–217.
- Kleber Neves and Olavo B. Amaral (2020). Addressing selective reporting of experiments through predefined exclusion criteria. eLife 9, e56626.
- Dan Lewer, Thomas D. Brothers, Elizabeth O'Nions and John Pickavance (2025). Factors associated with: Problems of using exploratory multivariable regression to identify causal risk factors. BMJ Medicine 4(1), e001375.
Last reviewed 2026-09-13.