Publication bias
Also known as file drawer problem
Publication bias is the tendency for studies to be written up and published, or not, depending on what they found. Studies with positive, striking or statistically significant results make it into journals; studies that found nothing tend to stay unpublished, in what the psychologist Robert Rosenthal called the file drawer.
The flaw is that the published record becomes a filtered sample of the research that was done, and the filter is the result itself. Anyone who reads only what was published, however carefully, sees the studies that worked and not the ones that didn’t. An effect can look reliable when many attempts failed, and a real but modest effect can look large. No individual study needs to be wrong for the whole literature to mislead.
Examples
The newsletter of wins
A company’s product team runs about 40 website experiments a year. When a test shows a clear improvement, it gets a write-up in the monthly newsletter; tests that show no difference are quietly closed. At year’s end a manager adds up the improvements reported in the newsletters, gets a total of “+38% conversions”, and is puzzled that actual sales barely moved.
Every write-up may be accurate, but the newsletter only ever received the winners. Some of those winners were flukes that happened to cross the line, and the tests that came out flat or negative, which would have pulled the total back down, were never counted. Adding up a record that was selected for success produces a picture of success.
The experiment nobody wrote up
A researcher runs a well-designed study testing whether a memory technique helps people learn vocabulary. It finds no difference. She decides it isn’t worth the months it would take to write up, since “journals don’t want null results”, and moves on to a project that looks more promising.
No journal rejected anything and nothing was hidden on purpose. The decision is understandable, and in one large study it was the main way the file drawer filled: among 221 social science experiments, Franco and colleagues found that null results were far less likely even to be written up (see Evidence). Multiplied across a field, many such reasonable choices leave the studies that did find an effect to speak for the technique.
The careful review
A team conducts a systematic review of every published trial of a new study-skills course, using a search method set in advance, and pools 18 trials. The small trials report large benefits; the few large trials report benefits near zero. The pooled estimate shows a moderate benefit, and the review recommends the course.
Nobody cherry-picked here: the review included every trial it could find. But if small trials that found nothing were less likely to be published, the literature it searched was already missing them. Small trials have noisy results, so the ones that got published tend to be the ones that came out large by chance. That lopsided pattern, where small studies show bigger effects than large ones, is what a funnel plot is designed to reveal (see How to respond). A careful method applied to a filtered literature inherits the filter.
Variants
- Not written up: authors don’t submit results they think are uninteresting. Franco and colleagues found this was where publication bias mainly arose in their sample.
- Not accepted: editors or reviewers are less keen on null results. Turner and colleagues couldn’t tell whether the bias they found came from authors, journals or both.
- Published with a positive spin: a study with a negative or unclear result is published in a way that emphasizes a favorable outcome.
- Selective outcome reporting: within a published study, the outcomes that came out significant are reported and others are left out. The Catalogue of Bias lists this as a related reporting bias.
- Selective citation: once published, null results are cited less often (see suppressed evidence).
Not the same error: cherry-picking. In Suppressed evidence, a person arguing for a conclusion leaves out evidence they could have included. In publication bias, the missing evidence is gone before any reader arrives, removed by many separate decisions about what to publish. A reader who faithfully reports everything in print can pass publication bias along without choosing anything. The two combine easily, since selective citation adds a second filter on top of the first.
When it isn’t an error
- When studies are selected on design, not results. Leaving out studies that were too small, unblinded or never completed is sound, provided the rule would apply the same way whatever the studies found.
- When the whole set of studies is known. Trial registries, regulatory files and reviews that track down unpublished results make the missing studies visible, so conclusions can rest on all of them.
- When publication was decided before the results were in. In a Registered Report, a journal accepts a study on the strength of its question and methods before the data are collected.
- When a selection of successes is labeled as one. “Case studies of projects that worked” doesn’t claim to show how often projects work.
The test: could this study’s result have affected whether you got to see it?
Looks like it, but isn’t
A small study turned down
A journal rejects a study of 24 people testing a new teaching method, explaining that a sample that small can’t reliably detect an effect of the size the authors were looking for. The journal applies the same rule to small studies that report positive results.
The study went unpublished, but not because of what it found: the reason is its size, and it would have been rejected just the same with a striking result. That’s selection on design, not results. Applying the rule consistently matters. If only the small null studies were turned down, it would be publication bias with a methodological excuse.
Trusting a positive Registered Report
A reader cites a published study that found a clear benefit. A colleague objects: “Of course it was positive. Only positive studies get published.” The reader points out that it was a Registered Report, accepted by the journal before any data were collected.
The colleague’s general worry is fair, but it doesn’t apply to this study. Its publication was guaranteed before its result existed, so a null finding would have appeared in the same journal. The study can still be wrong for other reasons; publication bias just isn’t one of them. That’s the decided before the results condition.
Why it happens
Researchers, reviewers and editors all find a clear result more interesting than “no difference”, and a null result is often ambiguous: the idea might be wrong, or the study might have been too small or badly run. So null results take effort to write up and persuade no one. The International Committee of Medical Journal Editors noted that trials showing a new treatment is inferior to standard treatment generate less enthusiasm, and that results that put financial interests at risk are “particularly likely to remain unpublished.” Funding, careers and attention all reward positive findings.
For readers, the problem is that missing studies leave no trace. A literature of positive findings looks complete unless you know what should be there, which is “what you see is all there is” applied to a whole field. Researchers who try many analyses and report the one that worked (P-hacking) add a second filter inside studies, on top of the filter between them. And because a field that publishes only its hits can make chance findings look like a pattern, the effect resembles the Texas sharpshooter fallacy at the scale of a literature. Judging a field by the studies that “survived” to publication also resembles Survivorship bias.
How to respond
- Ask whether you’re seeing all the studies. Is there a registry, a regulator’s file or a review that searched for unpublished results?
- Look for trial registration. Since 2005, journals in the International Committee of Medical Journal Editors have required clinical trials to be registered in a public registry at or before the start of enrollment, as a condition of publication. A registered trial can’t disappear without a trace, and its planned outcomes can be compared with the reported ones.
- Check for the funnel pattern. A funnel plot shows each study’s result against its size. Without bias, small studies scatter widely on both sides of the large ones. If the small studies all sit on the favorable side, some may be missing. Egger and colleagues, who proposed a statistical test for this, warn that its results should be “treated with considerable caution” when a meta-analysis rests on a few small trials, and the Catalogue of Bias notes that such plots need careful interpretation.
- Give extra weight to Registered Reports and large preregistered studies, whose publication didn’t depend on their results.
- Don’t conclude the effect is zero. Publication bias means the published estimate is probably too high, not that nothing is there. In the antidepressant trials Turner and colleagues examined, the effect shrank once unpublished results were counted but didn’t vanish.
Evidence
Publication bias is a flaw in how evidence reaches readers rather than an effect with a replication record, but its size has been measured directly by comparing published studies with complete cohorts of studies known to have been done.
- Rosenthal (1979) named it. In what he called the “extreme view” of the problem, “the journals are filled with the 5% of the studies that show Type I errors” (false positives), “while the file drawers back at the lab are filled with the 95% of the studies that show nonsignificant” results. As the paper’s title says, its concern was how much “tolerance for null results” a body of findings has.
- Turner and colleagues (2008) compared the U.S. Food and Drug Administration’s records of 74 trials of 12 antidepressants with the medical literature. Of the trials the FDA judged positive, 37 of 38 were published. Those it judged negative or questionable were, with three exceptions, either not published (22 trials) or published in a way that, in the authors’ opinion, conveyed a positive outcome (11 trials). The published literature made 94% of the trials look positive, against 51% in the FDA’s analysis, and the apparent effect size was 32% larger overall (11% to 69% for individual drugs).
- Franco, Malhotra and Simonovits (2014) followed 221 survey experiments run through a U.S. National Science Foundation program (Time-sharing Experiments for the Social Sciences), which had all passed peer review before being run. Studies with strong results were 40 percentage points more likely to be published than those with null results, and 60 points more likely to be written up at all. The bias arose mainly because authors didn’t write up and submit null findings.
- Schmucker and colleagues (2014) reviewed cohorts of studies approved by ethics committees or entered in trial registries. Only about half were ever published as journal articles, and approved studies with statistically significant results were more likely to be published (pooled odds ratio 2.8).
Remedies. After registration became expected, the pattern of results changed in at least one large set of trials. Kaplan and Irvin (2015) found that 17 of 30 large cardiovascular trials funded by the U.S. National Heart, Lung, and Blood Institute and published before 2000 (57%) showed a significant benefit on their main outcome, against 2 of 25 (8%) published afterward, and that preregistration on ClinicalTrials.gov was strongly associated with the shift. The design doesn’t show that registration caused it. Scheel, Schijen and Lakens (2021) found positive results for the first hypothesis in 96% of a random sample of standard psychology papers, against 44% of Registered Reports. They suggest reduced publication bias and fewer false positives as a plausible part of the explanation, among others they discuss.
What remains uncertain is how much publication bias changes conclusions in any particular field, which is harder to measure than how often it happens. The Catalogue of Bias notes that research has mostly documented its prevalence rather than its impact.
Sources
- Robert Rosenthal (1979). The file drawer problem and tolerance for null results. Psychological Bulletin 86(3), 638–641.
- Erick H. Turner, Annette M. Matthews, Eftihia Linardatos, Robert A. Tell and Robert Rosenthal (2008). Selective publication of antidepressant trials and its influence on apparent efficacy. New England Journal of Medicine 358(3), 252–260.
- Annie Franco, Neil Malhotra and Gabor Simonovits (2014). Publication bias in the social sciences: Unlocking the file drawer. Science 345(6203), 1502–1505.
- Christine Schmucker and colleagues (2014). Extent of non-publication in cohorts of studies approved by research ethics committees or included in trial registries. PLoS ONE 9(12), e114023.
- Matthias Egger, George Davey Smith, Martin Schneider and Christoph Minder (1997). Bias in meta-analysis detected by a simple, graphical test. BMJ 315(7109), 629–634.
- Catherine De Angelis and colleagues (International Committee of Medical Journal Editors) (2004). Clinical trial registration: A statement from the International Committee of Medical Journal Editors. New England Journal of Medicine 351(12), 1250–1251; also CMAJ 171(6), 606–607.
- Robert M. Kaplan and Veronica L. Irvin (2015). Likelihood of null effects of large NHLBI clinical trials has increased over time. PLoS ONE 10(8), e0132382.
- Anne M. Scheel, Mitchell R. M. J. Schijen and Daniël Lakens (2021). An excess of positive results: Comparing the standard psychology literature with Registered Reports. Advances in Methods and Practices in Psychological Science 4(2).
- Nicholas DeVito and Ben Goldacre (2019). Publication bias. Catalogue of Bias.
Last reviewed 2026-09-13.