Prosecutor's fallacy
Also known as fallacy of the transposed conditional or inverse fallacy
The prosecutor’s fallacy is confusing two probabilities that sound alike: how likely the evidence would be if a person were innocent, and how likely the person is to be innocent given the evidence. A forensic match that only 1 in 1,000 innocent people would show gets reported as a 1 in 1,000 chance that the person who matches is innocent.
The flaw is that the second number depends on something the first leaves out: how many innocent people could have matched. Evidence that is rare in any one innocent person can still turn up in dozens of them when thousands could have been involved, so a match can be rare and yet leave the person who matches more likely innocent than guilty. The name comes from the courtroom, but the same swap happens anywhere a probability of evidence gets read as a probability of a conclusion. It’s a formal fallacy in the sense that the problem is in the structure of the probability claim, not in what the evidence is about.
Examples
The fiber match
A fiber found at a warehouse break-in matches the jacket of a man whose only link to the case is that match. An analyst testifies that only 1 in 1,000 jackets would match. The prosecutor tells the jury: “That means there’s only a one-in-a-thousand chance he’s innocent.”
The 1 in 1,000 is the chance that an innocent person’s jacket would match. Suppose any of 50,000 adults in the city could have been there. About 50 of them would be expected to have matching jackets too, so the match puts him in a group of roughly 51 people who fit the evidence. On this evidence alone, the chance he’s the one is about 2%, not 99.9%. The match is still real evidence: it narrowed 50,000 people to about 51. The error is reporting that narrowing as near-certainty.
A significant result
A researcher tests whether a new study technique raises exam scores. The result is statistically significant, with p = 0.01, and her summary says: “There is only a 1% probability that the technique has no effect.”
A p-value is the probability of getting results at least this extreme if the technique has no effect. It isn’t the probability that the technique has no effect, given these results. That second number also depends on how plausible it was beforehand that the technique works, the same way the jacket example depends on how many people could have matched. This version looks respectable because it comes with a formal test. Fenton, Neil and Berger note that p-values are often misread in exactly this way. The statistician Jacob Cohen described this line of reasoning as appearing, at least implicitly, “in article after article in psychological journals”.
The gate that almost never beeps
A shop’s security gate beeps as a customer walks out. “These gates almost never go off for someone who hasn’t taken anything,” the guard says. “So you’ve almost certainly got something.”
No probability is stated, but the swap is the same. Say the gate beeps for 1 in 500 customers who have paid for everything and for every shoplifter, and 5,000 people pass through in a day, 2 of them shoplifters. The gate beeps for the 2 shoplifters and for about 10 innocent customers, so roughly 10 of every 12 beeps are false alarms. “Almost never goes off for the innocent” is true, and “this person is almost certainly guilty” is false. The fallacy doesn’t need numbers; it needs only the reversal.
Form
| Prosecutor’s fallacy | Modus tollens | |
|---|---|---|
| Premise | If P, then Q would be very unlikely | If P, then not Q |
| Premise | Q | Q |
| Conclusion | P is very unlikely | Not P |
| Valid? | No | Yes |
The fallacy borrows the shape of modus tollens, which is valid: if P rules Q out and Q happened, P is false. But “very unlikely” isn’t “ruled out”. Unlikely things happen when there are many chances for them to happen, which is exactly the fiber and gate examples. Jacob Cohen (1994) showed that making modus tollens probabilistic in this way makes it formally invalid, with a counterexample (credited to Pollard and Richardson) in which both premises are true and the conclusion is plainly false:
If a person is an American, then he is probably not a member of Congress. This person is a member of Congress. Therefore, he is probably not an American.
In probability notation, the error is treating P(evidence | innocent), “the probability of the evidence given innocence”, as if it were P(innocent | evidence). The two are linked by Bayes’ theorem, which brings in how common innocence and the evidence each are overall; they come out equal only when those are equally common. Fenton, Neil and Berger call the swap “the error of the transposed conditional”.
Its deductive cousin is Affirming the consequent. Both run a conditional backwards: affirming the consequent reads “if P, then Q” as “if Q, then P”, and the prosecutor’s fallacy reads “if innocent, this evidence is unlikely” as “given this evidence, innocence is unlikely”.
Variants
- The courtroom form: a random match probability for a fiber, blood type or DNA profile reported as the probability that the defendant is not the source.
- The p-value form: the probability of the data if there’s no effect, read as the probability that there’s no effect, as in the study example.
- The diagnostic form: a test’s accuracy in people who have a condition read as the chance that a person who tests positive has it. This is the positive-test example in Base rate neglect. Sidebotham and Dare, writing for anesthetists, describe it as confusing a test’s sensitivity with its positive predictive value.
- The defense attorney’s fallacy, named in the same 1987 paper as the prosecutor’s fallacy: the opposite overcorrection. “Fifty people in the city would match, so the match is irrelevant” dismisses evidence that has narrowed 50,000 possible sources to about 50, and that other evidence may narrow further.
When it isn’t an error
- When the right probability has actually been worked out. Saying “given the match, he’s very likely the source” is fine if it accounts for how many others could have matched and for the rest of the evidence.
- When the evidence is reported as a comparison. “This evidence is 1,000 times more likely if he’s the source than if he isn’t” compares two probabilities of the evidence and makes no claim about guilt. Fenton, Neil and Berger note that presenting evidence this way, as a likelihood ratio, avoids most common cases of the fallacy.
- When the group of possible sources is small. If only a handful of people could have produced the evidence, few innocent people are available to match by chance, and a rare match can be close to decisive.
- When the claim is only about the evidence. “An innocent person would rarely match” is true and relevant. The error is the swap, not the statistic.
The test: is this the probability of the evidence if the claim is false, or the probability that the claim is false given the evidence? If the second, where did the number of other possible matches come in?
Looks like it, but isn’t
Three vans in the lot
A delivery van scraped a parked car overnight in a gated lot. The gate log shows only three vans entered that night, and two of them left before anyone could look at them. Paint left on the car matches the van still parked there, and that color is found on about 1 in 1,000 vans. The lot manager concludes that the parked van almost certainly did it.
This has the surface of the fallacy: a rare match followed by near-certainty. But the manager isn’t reading the 1 in 1,000 as the chance of innocence. The gate log has already shrunk the possible sources to three, so the only way the parked van is innocent is if one of the other two vans also carries that rare color. With just two others, that chance is about 2 in 1,000, which leaves the parked van the source with a probability of about 99.8%. That’s the small group of possible sources condition.
A report that stops at the evidence
A forensic report on glass fragments found on a jacket says: “These findings are about 500 times more likely if the fragments came from the broken shop window than if they came from some other source.” It says nothing about how likely the jacket’s owner is to have broken the window.
Someone quick to cry fallacy might object that a probability is being used to point at a suspect. But both probabilities in the report are probabilities of the evidence, compared under two explanations. That’s the comparison condition. The fallacy would enter only if someone restated it as “500 to 1 that he broke the window”.
Why it happens
The two statements differ only in word order. “The chance of a match if he’s innocent” and “the chance he’s innocent if he matches” sound almost identical, and the first is the number that a forensic analyst can actually calculate. Eldridge’s review of juror comprehension quotes a clear illustration from Sjerps and Biesheuvel: “If I am a monkey, then it is highly likely that I have two eyes, two arms and two legs” doesn’t make “If I have two eyes, two arms and two legs, then it is highly likely that I am a monkey” true.
The missing ingredient is the one that Base rate neglect leaves out: how common each possibility is to begin with. A rare match feels like it speaks directly about the person in front of you, while the thousands of other people who could have matched are nowhere in view. The two errors overlap without being the same. Villejoubert and Mandel, who call the swap the inverse fallacy, found in a study of 45 people that how far their probability estimates strayed from the correct answers could be predicted by how far the inverse probabilities strayed from them, and discuss what separates the inverse fallacy from base rate neglect.
How the numbers are presented matters. In Thompson and Schumann’s (1987) experiments, students judging a written criminal case made more errors favoring the prosecution when a match statistic was given as a conditional probability, and more errors favoring the defense when it was given as a percentage. Most failed to spot the flaw in one or both of two fallacious arguments put to them. Compared with a Bayesian calculation, participants in both experiments also tended to give the statistical evidence too little weight overall, so the swap is not the only way people misjudge it. Eldridge credits their paper with coining the name “prosecutor’s fallacy”.
The error has reached real courtrooms. In the English case of Sally Clark, a mother convicted of murdering her two infant sons, an expert witness put the chance of two sudden infant deaths in a family like hers at 1 in 73 million. Fenton, Neil and Berger point out two problems with that figure: it wrongly assumed the two deaths were independent, and it was presented in a way that may have led the jury into the prosecutor’s fallacy, reading the rarity of two such deaths as the rarity of innocence. The relevant comparison was with the alternative explanation, a double murder, which is also very rare. Her convictions were quashed in 2003; Sidebotham and Dare note that by then the appeal also had microbiological findings that had not been presented at her trial.
How to respond
- Say both probabilities out loud, in full. “The chance of this evidence if he’s innocent” and “the chance he’s innocent given this evidence” are different questions. Saying them side by side usually makes the swap visible.
- Count the other possible matches. Multiply the number of people (or cases) that could have produced the evidence by the chance that any one of them would. If that comes to more than a few, the evidence alone can’t make one match near-certain.
- Ask for the comparison, not just the rarity. How much more likely is the evidence under one explanation than the other? That tells you how strong the evidence is without pretending to settle the conclusion.
- Use counts instead of percentages. Sorting 1,000 or 50,000 imaginary people into those who match and those who don’t builds the missing piece into the picture. The research on this remedy, which helps substantially without solving the problem, is described in Base rate neglect.
Sources
- William C. Thompson and Edward L. Schumann (1987). Interpretation of statistical evidence in criminal trials: The prosecutor's fallacy and the defense attorney's fallacy. Law and Human Behavior 11(3), 167–187.
- Jacob Cohen (1994). The earth is round (p < .05). American Psychologist 49(12), 997–1003.
- Gaëlle Villejoubert and David R. Mandel (2002). The inverse fallacy: An account of deviations from Bayes's theorem and the additivity principle. Memory & Cognition 30(2), 171–178.
- Norman Fenton, Martin Neil and Daniel Berger (2016). Bayes and the law. Annual Review of Statistics and Its Application 3, 51–77.
- Heidi Eldridge (2019). Juror comprehension of forensic expert testimony: A literature review and gap analysis. Forensic Science International: Synergy 1, 24–34.
- David Sidebotham and Tim Dare (2025). Flipping the conditional: Why we are probably wrong about probabilities. Pediatric Anesthesia 35(8), 584–589.
Last reviewed 2026-09-13.