Observer-expectancy effect
Also known as experimenter expectancy effect or experimenter bias
The observer-expectancy effect is when a researcher’s expectations about a study’s outcome influence the outcome. It works in two ways. The researcher may see and record what they expect, especially when a measurement calls for judgment. Or they may change what happens, treating subjects they expect to do well a little differently or giving off cues that the subjects pick up, without realizing it.
The flaw is that the measurement is no longer independent of the hypothesis it’s testing. A result that seems to confirm an expectation was partly produced by that expectation, so it can’t count as evidence for it. The researcher can be entirely honest and unaware of any influence, which is exactly why it’s hard to catch without a design that prevents it.
Examples
The horse that could count
In Germany in the early 1900s, a horse called Clever Hans answers arithmetic questions by tapping his hoof, and gets them right. A 1904 commission of experts finds no sign of deliberate tricks, and Hans answers correctly even when strangers question him in his owner’s absence. Many conclude he can really count.
The psychologist Oskar Pfungst found that Hans got the answers right only when the person asking knew them. When they did, he was right about 98% of the time in one series of tests; when nobody present knew the answer, 8%. Questioners leaned forward slightly as Hans began tapping and gave a tiny upward jerk of the head when he reached the right number, and Hans stopped. The movements were involuntary and so slight that questioners didn’t notice making them; Pfungst found he made them himself. The expectation of the right answer, carried by those movements, was producing the right answer.
The therapist who rates the progress
A clinic compares a new exercise program for knee pain with its standard one. The physical therapists who deliver the programs also score each patient’s walking at the end, on a 1-to-10 scale of “gait quality”. Patients in the new program score higher, and the clinic adopts it.
The therapists knew which program each patient was in, and they hoped the new one would work. A subjective rating leaves room for that knowledge to tip borderline judgments, and it doesn’t take many: in a review of clinical trials whose yes-or-no outcomes were assessed by both blinded and unblinded assessors, the exaggerated results came from reclassifying a median of only 3% of patients per trial (see Evidence). The program may work, but this comparison can’t separate its effect from the raters’ hopes.
The encouraging assistant
A research assistant runs a study on whether a short pep talk improves performance on a timed puzzle. She reads the pep talk to half the participants. She’s careful to time everyone identically, but when the pep-talk group is working, she smiles and says “you’re doing great” a little more often.
Here nothing is misrecorded: the stopwatch is accurate. The expectation acts earlier, on the participants themselves. If the pep-talk group does better, some of that may come from the extra encouragement, which is a second treatment nobody planned. This is the kind of influence Robert Rosenthal studied under the name interpersonal expectancy effects.
Variants
- Observer bias: expectations affect what is recorded or how it’s interpreted, as with the knee ratings. The Catalogue of Bias uses “observer bias” more broadly, for any systematic difference between a true value and the value observed, including habits like rounding blood pressure readings.
- Influencing the subjects: expectations change how the researcher treats the subjects, or give off cues they respond to, as with Clever Hans and the pep talk. Holman and colleagues give examples such as feeding control animals differently or giving patients different nonverbal cues.
- Expectations outside research: Rosenthal extended the idea to teachers, employers and therapists. The best-known case is contested: a classroom study by Rosenthal and Lenore Jacobson claimed that pupils whose teachers were told they were “ready to bloom”, although they had been picked at random, did better on an IQ test at the end of the year. Reviewing 35 years of research, Jussim and Harber (2005) concluded that such self-fulfilling prophecies do occur but are typically small and may fade rather than build up, and that teachers’ expectations may predict students’ outcomes more because they are accurate than because they are self-fulfilling.
Not the same: demand characteristics. In the observer-expectancy effect, the researcher’s expectations change the result. Demand characteristics, a concept from the psychologist Martin Orne, work from the other side: participants work out what the researcher is investigating or hopes to find, and adjust their behavior accordingly, for example by trying to be a “good subject”. The two can combine, since a researcher’s cues are one way participants learn what’s expected. Similarly, a placebo effect comes from participants’ expectations about a treatment, not the researcher’s.
When it isn’t an error
- Whoever records the outcome doesn’t know what to expect for each subject: they’re blind to which group the subject is in or to the hypothesis.
- Whoever deals with the subjects doesn’t know either, or the procedure is standardized or automated so there’s little room for different treatment.
- The outcome leaves no room for judgment. The Catalogue of Bias notes that objective data, such as death, are at much lower risk of observer bias than subjective judgments.
- An expectation is accurate and was checked independently. Expecting a result doesn’t create a bias if the result is measured by someone or something the expectation can’t reach.
The test: could the person recording or handling this subject have known, and been influenced by, what result was expected?
Looks like it, but isn’t
An open trial with a blinded assessor
In a trial of a new exercise program, patients and their therapists obviously know who is doing which program. At the end, a separate assessor, who never learns anyone’s group, times each patient over a fixed walking course with an electronic timer.
A critic might note that the trial “wasn’t blinded”. But the people who knew the groups didn’t measure the outcome, and the measurement leaves little room for judgment. The patients’ own expectations could still matter, but the researchers’ expectations have no route into the recorded result. That’s the blind recorder and no room for judgment conditions.
The trainer who picks the winners
An experienced guide-dog trainer, watching a litter of puppies at eight weeks, predicts which will pass the program. She has no further contact with them. A year later, instructors who were never told her predictions assess the dogs, and most of her picks pass.
Her expectations matched the result, which can look like expectancy at work. But she couldn’t influence the dogs’ training or their assessment: the instructors didn’t know what she expected. Her predictions came true because she read real signs in the puppies. That’s the accurate expectation, checked independently condition, and it’s the alternative Jussim and Harber raise for teachers.
Why it happens
The influence is usually invisible to the person exerting it. Pfungst, after identifying the cues, played the horse’s role himself and tested 25 people of different ages and backgrounds, who concentrated on a number while he tapped. All but two made the same involuntary movements, and only in a few isolated instances did anyone report being aware of moving. He found he had been making the cues himself when questioning Hans.
Recording is vulnerable for a related reason. Holman and colleagues describe observer bias as a consequence of confirmation bias: people tend to notice and remember outcomes that fit what they already believe, and a researcher who expects a result has that belief. The biases are strongest, they write, when researchers expect a particular result, when the variable is subjective and when there’s an incentive to confirm predictions.
The effect has a family resemblance to the Texas sharpshooter fallacy, where a forensic analyst’s knowledge of a suspect can shift what counts as a match, and to P-hacking, where a researcher’s hopes shape the analysis rather than the data collection. Participants can bring a parallel distortion to a study through their memories, as in Recall bias.
How to respond
- Ask who knew. Did the people measuring the outcome, or dealing with the subjects, know which group each subject was in or what was expected?
- Prefer blinding. The Catalogue of Bias names blinding outcome assessors to participants’ exposure as a key method, including separating data on exposures from data on outcomes. Where the people delivering a treatment can’t be blinded, a separate blinded assessor often can be.
- Prefer measurements with little room for judgment, or automated recording.
- Don’t rely on training or good intentions alone. The Catalogue of Bias describes a study in which training reduced differences between nurses measuring blood pressure but didn’t remove them, and concludes that observer bias can be reduced but will likely always remain.
- Don’t dismiss the result. An unblinded study is weaker evidence, not proof that its finding is wrong.
Evidence
The observer-expectancy effect is a flaw in method rather than an effect with a single replication record, but its existence is well documented and its typical size has been estimated. How large and how common it is has been disputed from the start.
- Clever Hans (Pfungst, 1911). With the questioner knowing the answer, 90% to 100% of Hans’s responses in various test series were correct; without, 10% at most. With blinders that stopped him seeing the questioner, he failed.
- Rosenthal and Fode (1963), as described by Holman and colleagues and by Bishop: students who were told their rats had been bred to be good or poor at mazes recorded better performance for the “bright” rats, although the rats were ordinary and randomly assigned. Bishop describes the result as suggesting the students tried harder with animals they thought had potential; whether handling or recording explains it isn’t settled by these accounts.
- Rosenthal and Rubin (1978) summarized 345 experiments on interpersonal expectancy effects, in areas ranging from reaction time and animal learning to person perception and everyday life, as an early quantitative summary of a whole field.
- The early critique. Bishop (2020) points out that T. X. Barber and M. J. Silver (1968) criticized this research for poor methodological quality, including what Bishop calls p-hacking, and concluded that experimenter bias effects, while real, were far less common and smaller than Rosenthal’s early work implied.
- Hróbjartsson and colleagues (2012) reviewed 21 randomized clinical trials (4,391 patients) in which the same binary outcome, mostly subjective, was assessed by both blinded and unblinded assessors. Unblinded assessors exaggerated the treatment effect (the odds ratio) by 36% on average, even though the two kinds of assessor agreed in a median of 78% of assessments. The same group’s later reviews, cited by the Catalogue of Bias, found exaggerated effects with unblinded assessment for measurement-scale outcomes and time-to-event outcomes too.
- Holman and colleagues (2015) compared 83 pairs of closely matched evolutionary biology studies, one done blind and one not. On average the nonblind studies reported effects about 27% larger, and in a text-mining study of life science papers, nonblind papers reported more p-values below .05. They caution that the data are nonexperimental and note counterexamples in which blinding made no significant difference, but conclude that observer bias is probably an important cause.
What remains uncertain is how large the effect is outside settings with subjective outcomes, and how much the early laboratory findings overstated it.
Sources
- Oskar Pfungst (translated by Carl L. Rahn) (1911). Clever Hans (the Horse of Mr. von Osten): A Contribution to Experimental Animal and Human Psychology. Henry Holt and Company (Project Gutenberg eBook 33936).
- Robert Rosenthal and Donald B. Rubin (1978). Interpersonal expectancy effects: The first 345 studies. Behavioral and Brain Sciences 1(3), 377–386.
- Dorothy V. M. Bishop (2020). The psychology of experimental psychologists: Overcoming cognitive constraints to improve research. Quarterly Journal of Experimental Psychology 73(1), 1–19.
- A. Hróbjartsson and colleagues (2012). Observer bias in randomised clinical trials with binary outcomes: Systematic review of trials with both blinded and non-blinded outcome assessors. BMJ 344, e1119.
- Luke Holman, Megan L. Head, Robert Lanfear and Michael D. Jennions (2015). Evidence of experimental bias in the life sciences: Why we need blind data recording. PLOS Biology 13(7), e1002190.
- Lee Jussim and Kent D. Harber (2005). Teacher expectations and self-fulfilling prophecies: Knowns and unknowns, resolved and unresolved controversies. Personality and Social Psychology Review 9(2), 131–155.
- Jim McCambridge, Marijn de Bruin and John Witton (2012). The effects of demand characteristics on research participant behaviours in non-laboratory settings: A systematic review. PLoS ONE 7(6), e39116.
- K. Mahtani, E. A. Spencer and J. Brassey (2017). Observer bias. Catalogue of Bias.
Last reviewed 2026-09-13.