Outcome bias
Outcome bias is judging a decision by how it turned out instead of by the information and reasoning available when it was made. The same choice, made with the same knowledge, gets called smart when luck goes its way and careless when it doesn’t.
The flaw is that a result can’t reach back and change the decision. Where chance plays a part, a good decision can end badly and a bad one can end well. What the decision maker knew, what the odds were and what the options would cost are the same whichever way it came out, so those are what a verdict on the decision has to rest on. The outcome tells you about the world; on its own, it tells you little about the choice.
Examples
The fourth-down call
A youth football coach, facing fourth-and-two near midfield in the second quarter, goes for it instead of punting. The play is stopped a yard short, the other team scores, and the team goes on to lose. A parent in the stands says: “What a reckless call. You punt there.” Last month the same coach made the same call in the same spot, it worked, and the same parent called it gutsy.
The situation, the odds and the coach’s reasoning were the same both times. Only the result differed, and the verdict flipped with it. This is the core case: identical decisions rated differently because of how they came out.
“See, no need to leave early”
A colleague leaves for the airport 45 minutes before boarding on a Friday afternoon, skipping the margin everyone else builds in. Traffic happens to be light and they make the flight. Back at the office they say: “Told you. All that extra time is a waste.” The team starts planning trips the same way.
This is the version that runs in the flattering direction, and it is easier to miss because nothing went wrong. The plan relied on luck, and the luck held. Treating the success as proof the plan was sound rewards the risk and invites a repeat, the kind of strategy revision after narrow results that researchers have found in professional coaches (see Evidence).
Two teams, one checklist
Two teams release software updates the same week, following the same review checklist. One update goes out cleanly. The other hits a rare hardware fault at the data center and the service is down for an hour. In the performance review, the second team’s lead is marked down for “poor judgment on release timing.”
Nothing about the second team’s process differed; the fault was outside what either team could have checked. A formal review can look like careful evaluation while grading luck. Experiments in which people pay someone to make risky investments on their behalf find the same thing: they reward the agent more after a lucky result even when they know exactly what the agent did.
When it isn’t an error
- When you don’t know what the decision maker knew. If you can’t see their information or reasoning, how things turned out is some evidence about both. It’s weak evidence from a single case, but not irrelevant.
- When decisions repeat. Across many similar decisions, results add up to a track record, and a track record does reveal whether a strategy or a person’s judgment is any good. Luck evens out; a single outcome doesn’t.
- When the outcome exposes a mistake that was knowable beforehand. If a bad result leads you to discover that someone ignored a warning or misread the odds, you’re judging what they should have known at the time. The outcome only prompted the review.
- When you’re deliberately rewarding results, not judging the decision. Paying for results can be a sensible way to set incentives, as long as nobody mistakes the payout for a verdict on the quality of the choice.
The test: if the result had gone the other way, with nothing else changed, would you rate the decision differently? If so, you’re rating the outcome.
Looks like it, but isn’t
The staffing agency
A café owner has hired through the same staffing agency ten times. Seven of those hires quit within a month. “The agency’s candidates don’t work out for us. I’m going to find staff another way.”
The owner is judging by outcomes, but not by one outcome. Ten similar decisions with a consistent result are a track record, and that track record is real evidence about the agency. That’s the decisions repeat condition above.
The skipped inspection
A homeowner’s new deck sags within a year. The inspector finds the builder skipped the footing depth that the local code requires. The homeowner says: “Building it that way was a bad decision.”
The sagging deck is what prompted the review, but the verdict rests on something known before any board was laid: the code requirement and the reason for it. A builder who followed the code and still had a deck fail for an unforeseeable reason would deserve a different verdict. That’s the mistake knowable beforehand condition.
Why it happens
Jonathan Baron and John Hershey suggested that knowing the outcome changes which arguments come to mind. After a failure, the reasons against the choice stand out; after a success, the reasons for it do. Their participants said outcomes shouldn’t affect their ratings of a decision, and rated by outcome anyway.
Part of the pull is that outcomes often are informative. Where good results tend to follow good decisions, using results as a guide is a reasonable habit that misfires when chance plays a large part or when the decision maker’s information is fully known.
It is easy to confuse with Hindsight bias, and the two often travel together. Hindsight bias distorts your sense of how predictable the outcome was: after a bad result, it seems the risk was obviously high. Outcome bias is about how you grade the decision, and it appears even when the odds are stated and not in dispute. In Baron and Hershey’s operation scenario, participants were told that 8% of patients die from the procedure; the failure didn’t change that number, yet it still lowered their rating of the decision. In moral judgments the two can chain together: Markus Kneer and Izabela Skoczeń found that a harmful outcome made people rate the harm as having been more probable, which in turn made them judge the person more negligent.
A related pattern concerns your own results. The Self-serving bias is taking credit for successes and blaming failures on luck. Outcome bias would have you judge your own failed decision as a bad one; self-serving bias would have you blame the dice. The two can pull in opposite directions, which is one reason people can seem harsh about others’ bad luck and forgiving of their own. When a result colors judgments of a person’s other qualities, not just the decision, it shades into a Halo effect.
How to respond
- Judge the decision before looking at the result, when you can. Ask what the person knew, what the options and odds looked like and what a reasonable person would have done, and settle that first.
- Run the flip. Imagine the coin had landed the other way, with everything else the same. If your verdict changes, it was about the outcome.
- Look for a track record. One result says little about skill. Several similar decisions say more.
- Record reasoning at the time. A note of the options and odds considered gives a later review something other than the result to judge. This is sensible practice rather than a tested remedy.
Whether warnings or instructions reduce outcome bias is not well established. Knowing the principle isn’t enough: in both Baron and Hershey’s study and the 2023 replication, people who said outcomes shouldn’t count showed the bias anyway. Kneer and Skoczeń tested three strategies against the hindsight side of the problem in blame judgments and found some, but not all, promising.
Evidence
Status: replicates robustly. A preregistered replication with nearly 700 participants found the original effect, larger than first reported, and experiments with real money and field data from professional sports point the same way. One preregistered replication of a related claim about moral judgments did not succeed.
- Baron and Hershey (1988) ran five studies with undergraduates, who evaluated decisions made under uncertainty about medical treatments or monetary gambles, knowing they had all the information the decision maker had had. They rated the thinking as better, the decision maker as more competent, and were more willing to let that person decide for them when the outcome was good. In their first experiment, 20 participants each rated 15 medical cases. The one later used in the replication below describes a man with a heart condition deciding on a bypass operation that would relieve his pain but that 8% of patients don’t survive.
- Aiyer and colleagues (2023) preregistered a replication of that first experiment with 692 online participants, changing it so that each person saw only one version (success or failure) rather than all of them. The same operation decision was rated clearly worse when it failed (effects of d = 0.77 to 1.1, against d = 0.21 to 0.53 in the original). (Here d measures the size of a difference between groups; 0.2 is conventionally “small”, 0.5 “medium” and 0.8 “large”.) Participants who said outcomes should not be taken into account still showed the bias (d = 0.64). The authors note a limit: it was a single scenario.
- König-Kersting and colleagues (2021) ran three incentivized experiments in which one person made risky investments for another. The investor’s evaluations of, and payments to, the agent depended strongly on the random result, even though they knew exactly what the agent had chosen.
- Lefgren, Platt and Price (2015) studied professional basketball coaches and found they were more likely to change strategy after a loss than a win, even after narrow losses that say almost nothing about how good a team is, and even when a loss had been expected.
- Kneer and Skoczeń (2023), in ten preregistered experiments with 2,043 participants, found that harmful outcomes raised people’s estimates of how probable the harm had been and, through that, their judgments of negligence and blame.
- A failed replication in moral judgment. Gino, Shu and Bazerman (2010) had reported that people judge an unfair choice as less ethical when chance makes it harmful. A preregistered replication by Li and colleagues (2025, 236 participants, posted as a preprint and not peer reviewed) found a difference in the same direction but much smaller and not statistically significant.
What remains uncertain is how large the bias is outside laboratory scenarios, and how much of it runs through hindsight, a distorted sense of how likely the outcome was, rather than through the evaluation of the decision itself. Its reach into judgments of ethics, as opposed to judgments of decision quality, is less clear. That outcomes sway judgments of decisions made with identical information is well established.
Sources
- Jonathan Baron and John C. Hershey (1988). Outcome bias in decision evaluation. Journal of Personality and Social Psychology 54(4), 569–579.
- Lars Lefgren, Brennan Platt and Joseph Price (2015). Sticking with what (barely) worked: A test of outcome bias. Management Science 61(5), 1121–1136.
- Christian König-Kersting, Monique Pollmann, Jan Potters and Stefan T. Trautmann (2021). Good decision vs. good results: Outcome bias in the evaluation of financial agents. Theory and Decision 90(1), 31–61.
- Sriraj Aiyer, Hoi Ching Kam, Ka Yuk Ng, Nathaniel A. Young, Jiaxin Shi and Gilad Feldman (2023). Outcomes affect evaluations of decision quality: Replication and extensions of Baron and Hershey's (1988) outcome bias Experiment 1. International Review of Social Psychology 36(1), article 12.
- Markus Kneer and Izabela Skoczeń (2023). Outcome effects, moral luck and the hindsight bias. Cognition 232, article 105258.
- Sophia Li, Kelly Hu, Don A. Moore and Max H. Bazerman (2025). Replication of Gino, Shu, and Bazerman (2010, Study 2). PsyArXiv preprint (not peer reviewed).
Last reviewed 2026-09-13.