P-hacking
Also known as data dredging or fishing expedition
P-hacking is trying different ways of handling or analyzing a set of data (dropping some observations, switching the outcome measure, adding a control variable, collecting a few more participants) until a result crosses the threshold for statistical significance, usually p < .05, and then reporting that version as if it were the analysis that had been planned all along.
The flaw is that the threshold only means what it claims if you get one try. A 5% threshold promises that when there’s no real effect, a test will wrongly come out significant about one time in twenty. Every extra way of analyzing the same data is another try, so the chance that at least one of them comes out significant by luck climbs well above 5%. Reporting the try that worked, without the others, presents the result as far stronger evidence than it is. None of the individual choices has to be unreasonable, and the person making them usually isn’t trying to deceive anyone.
Examples
The analysis that got there in the end
A graduate student tests whether a pleasant scent in the room helps people solve puzzles. The first analysis gives p = .09. Two participants took much longer than everyone else, so she excludes them as outliers: p = .06. Adding participants’ age as a control variable gives p = .04. The write-up reports the final analysis: “The scent significantly improved puzzle performance (p = .04).”
The clear-cut case. Excluding outliers and controlling for age can each be defensible, but here each step was taken because the previous result fell short, and the report doesn’t mention the earlier versions. Had the first analysis given p = .04, she would have stopped there. The .04 is the result of three attempts, reported as if it were one.
Checking every morning
A website team runs an experiment comparing two sign-up buttons. They planned to run it for a month, but someone checks the results every morning. On day 9 the new button’s advantage is significant, so they stop the test and announce the winner.
Nothing was excluded and only one comparison was made, which is why this looks like diligence rather than p-hacking. But a running tally of random data drifts above and below the significance line, and checking it daily gives it many chances to cross. Stopping the first time it does turns the drift into a “finding”. Simmons, Nelson and Simonsohn showed in simulations that even a single extra look (testing at 20 participants per group and, if that fails, again at 30) raised the false-positive rate by about half.
One analysis, chosen after looking
A researcher predicts that a short breathing exercise will reduce test anxiety. Once the data are in, she sees that the anxiety questionnaire is noisy but that heart rate shows a clear difference, so she makes heart rate her main measure. She runs exactly one test, which is significant, and reports it.
This is the boundary case. She tried only one analysis and had her hypothesis before collecting data, so it isn’t p-hacking in the strict sense of trying test after test. Andrew Gelman and Eric Loken call it the garden of forking paths and argue it causes the same problem: if the data had looked different, she would reasonably have chosen a different measure, so the one test she ran was still picked from many she could have run, and the p-value doesn’t account for that. The problem, as they put it, can arise “even when there is no ‘fishing expedition’ or ‘p-hacking’ and the research hypothesis was posited ahead of time.”
Variants
Simmons and colleagues call the choices involved researcher degrees of freedom. Common ones:
- Flexible exclusions: deciding which observations count as outliers after seeing how each rule affects the result.
- Choosing among outcome measures: collecting several and reporting the one that worked.
- Optional stopping: checking results as data come in and stopping when they’re significant, or collecting more when they aren’t.
- Flexible control variables: adding or dropping covariates until the effect crosses the line.
- Dropping or combining conditions: reporting only the comparisons between groups that came out significant.
- The garden of forking paths: a single analysis whose details were chosen after seeing the data, with no conscious search at all. Gelman and Loken distinguish it from p-hacking proper, but the effect on the p-value is the same kind.
Usage varies. Lewer and colleagues treat “data dredging” and “fishing expeditions” as other names for p-hacking; the Catalogue of Bias uses data-dredging bias as a broader heading that includes p-hacking, fishing for the best statistical model, and HARKing.
Not the same error: the Texas sharpshooter fallacy. In the Texas sharpshooter fallacy, the target is drawn after the shots: a hypothesis or pattern is chosen to fit wherever the data happened to cluster, as in hunting through subgroups or HARKing (presenting a hypothesis formed from the results as if it had been set out beforehand). In p-hacking the target can stay where it was, a hypothesis named in advance, while the analysis is adjusted until the data hit it. The two often travel together, and writers use the terms loosely. Lewer and colleagues (2025) list p-hacking and the Texas sharpshooter fallacy as separate items among the poor research practices that one common study design combines, and treat the sharpshooter as another name for HARKing. Because sources treat the two as neighbors rather than one as a type of the other, this site doesn’t classify p-hacking as a kind of the Texas sharpshooter fallacy.
When it isn’t an error
Analyzing data in more than one way is normal. It’s sound when:
- The analysis was fixed before the data were seen, as in a preregistration, where the research questions and analysis plan are recorded before observing the outcomes. Departures from the plan are reported as departures.
- Exploration is labeled as exploration. Looking at data from many angles is how hypotheses are found. Gelman and Loken call refining hypotheses in light of the data “good scientific practice”; the problem is presenting the result as a confirmed test.
- All the analyses are reported. If a paper shows the result under each reasonable outlier rule, covariate set or measure, readers can see how much it depends on the choice.
- The number of tests is accounted for. Corrections for multiple comparisons, or a stricter threshold set in advance, build the extra tries into the evidence.
- A pattern found by exploring is confirmed in new data, analyzed by a plan fixed beforehand.
The test: would the analysis have been the same if the data had come out differently, and does the report show every version that was tried?
Looks like it, but isn’t
An exploratory finding, labeled
A team studying sleep and memory notices, in data collected for another purpose, that people who napped after lunch did better on an afternoon recall test. Their paper reports it in a section headed “Exploratory analyses”, says that it wasn’t predicted and that several other comparisons were examined and showed nothing, and calls for a preregistered test.
The team did exactly what p-hacking does, looking through the data for something significant. What makes it legitimate is that they said so: the reader knows it was one of many looks, knows it wasn’t planned, and is told it needs confirming. That’s the labeled exploration condition.
Every specification on the table
A study of whether a reading app improves test scores reports its preregistered analysis, then a table showing the result under three other outlier rules and with and without controls for prior grades. The effect is significant in four of the six versions, and the paper says so.
The researchers ran several analyses, and a skeptic might suspect they picked the best. But the planned analysis came first and every version is reported, including the ones that fell short. Readers can see the result is fairly, but not completely, robust to the choices, and the paper doesn’t hide which way the choices cut.
Why it happens
Simmons, Nelson and Simonsohn argue that p-hacking isn’t mainly a matter of bad intent but of two things combined: genuine ambiguity about how best to analyze data, and a researcher’s wish for a significant result. Faced with several defensible options, people tend to conclude, “with convincing self-justification,” that the right choice is the one that gives the answer they hoped for. As an example of the ambiguity, in about 30 articles from a single journal they found that the rules for excluding reaction times as “too fast” or “too slow” varied enormously. This is confirmation bias operating on analytic choices.
Once the results are known, the path that led to them also feels inevitable. Nosek and colleagues point to Hindsight bias as one reason researchers find it hard to keep ideas generated from data separate from predictions tested on it. And journals’ preference for significant results (Publication bias) gives the whole process its incentive: a study that stops at p = .09 is much harder to publish than one that finds its way to .04.
How to respond
- Ask whether the analysis was planned. Look for a preregistration, and compare it with what was reported: the outcome measures, exclusions, sample size and main test.
- Look for the other versions. Simmons and colleagues proposed that authors state their stopping rule in advance, list every variable and condition, and report results with and without any exclusions or covariates. When those are missing, ask what the result looks like without the choices that got it over the line.
- Be wary of a result that depends on one particular choice, especially in a small study with a p-value just under .05.
- Look for replication in new data. A result produced by flexible analysis often fails to reappear when the analysis is fixed in advance.
- Don’t treat preregistration as a cure-all. Gelman and Loken, who think it “might make sense” where new data are easy to gather, write that it “cannot realistically be a general solution”, for example in fields that reanalyze the same public data repeatedly. Disclosure of every choice matters either way.
- Don’t conclude the effect is false. A p-hacked result is weak evidence, not evidence of no effect.
Evidence
P-hacking is a flaw in method rather than an effect with a replication record, but both how much it distorts results and how often it happens have been studied.
How much it inflates false positives. Simmons, Nelson and Simonsohn (2011) simulated experiments with no real effect (20 observations per group) and counted how often at least one of a set of analyses came out significant at p < .05:
| Researcher degree of freedom | False-positive rate |
|---|---|
| None (a single planned test) | 5% |
| Two correlated outcome measures, or their average | 9.5% |
| Adding 10 more observations per group if the first test fails | 7.7% |
| Controlling for gender, or its interaction with the treatment | 11.7% |
| Dropping, or not dropping, one of three conditions | 12.6% |
| All four combined | 60.7% |
They also ran two real experiments to show how it works in practice. In the second, as first written up, 20 undergraduates listened to either the Beatles’ “When I’m Sixty-Four” or a control song, and afterward the first group was, according to their birth dates, “nearly a year-and-a-half younger” (p = .040), a result that is necessarily false. It came out significant only after controlling for participants’ fathers’ age (without it, p = .33), with a sample size not fixed in advance and results checked about every 10 participants, and with other measures and a third song condition collected but left out of the report. Their fully disclosed version of the report shows that 34 students had taken part.
How common it is. John, Loewenstein and Prelec (2012) surveyed more than 2,000 psychologists anonymously, with incentives for honest answers, and found that the share who admitted to questionable research practices was “surprisingly high”, suggesting that some may be the prevailing norm. Head and colleagues (2015) text-mined p-values across the sciences and found patterns consistent with widespread p-hacking, but concluded that its effect seemed weak relative to the real effects being measured and “probably does not drastically alter scientific consensuses drawn from meta-analyses”. The Catalogue of Bias notes that there is no definitive account of how often it happens or how severe its bias is, and that the distortion is worst when effects are small, measures imprecise, designs flexible and samples small.
What remains uncertain is how much of the published literature in any given field is affected, and how far preregistration and disclosure requirements reduce it in practice.
Sources
- Joseph P. Simmons, Leif D. Nelson and Uri Simonsohn (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science 22(11), 1359–1366.
- Andrew Gelman and Eric Loken (2013). The garden of forking paths: Why multiple comparisons can be a problem, even when there is no "fishing expedition" or "p-hacking" and the research hypothesis was posited ahead of time. Unpublished manuscript, Columbia University (14 November 2013).
- Andrew Gelman and Eric Loken (2014). The statistical crisis in science. American Scientist 102(6), 460–465.
- Leslie K. John, George Loewenstein and Drazen Prelec (2012). Measuring the prevalence of questionable research practices with incentives for truth telling. Psychological Science 23(5), 524–532.
- Megan L. Head, Luke Holman, Rob Lanfear, Andrew T. Kahn and Michael D. Jennions (2015). The extent and consequences of p-hacking in science. PLOS Biology 13(3), e1002106.
- Brian A. Nosek, Charles R. Ebersole, Alexander C. DeHaven and David T. Mellor (2018). The preregistration revolution. Proceedings of the National Academy of Sciences 115(11), 2600–2606.
- Norbert L. Kerr (1998). HARKing: Hypothesizing after the results are known. Personality and Social Psychology Review 2(3), 196–217.
- Dan Lewer, Thomas Brothers, Elizabeth O'Nions and John Pickavance (2025). Factors associated with: Problems of using exploratory multivariable regression to identify causal risk factors. BMJ Medicine 4(1), e001375.
- A. Erasmus, B. Holman and J. P. A. Ioannidis (2020). Data-dredging bias. Catalogue of Bias.
Last reviewed 2026-09-13.