Stereotype threat is the idea that people can score below their own usual level on a hard test when they are aware of a negative stereotype about a group they belong to — and the evidence for it is genuine, contested, and much smaller than its fame suggests. In settings built to resemble a real exam, the pooled effect is around a seventh of a standard deviation, meta-analyses find clear signs of publication bias, and a large pre-registered replication found nothing at all.
Where the idea came from
In 1995 Claude Steele and Joshua Aronson published a set of experiments in the Journal of Personality and Social Psychology under the title “Stereotype threat and the intellectual test performance of African Americans”. Their proposal was that in a situation where a negative stereotype about your group could apply to you, the awareness of it becomes an extra burden — and that this burden, rather than ability, can pull performance down on a difficult test.
Two features of the idea explain why it spread so far. It is situational, so it points at something that can be changed rather than at the people taking the test. And it does not require anyone to believe the stereotype: knowing that it exists and might be applied to you is enough. Within a decade the term had moved out of social psychology and into teaching handbooks, admissions policy and general conversation.
The claim the original study did not make
This is the part almost every popular account gets wrong, and it matters more than any other sentence on this page.
In 2004, Paul Sackett, Chaitra Hardison and Michael Cullen published a paper in American Psychologist documenting that the 1995 work was, in their words, “widely misinterpreted in both popular and scholarly publications as showing that eliminating stereotype threat eliminates the African American–White difference in test performance”. It did not show that. The scores in the experiments were statistically adjusted for students' prior SAT performance, and so what the results actually demonstrated was that in the absence of stereotype threat the two groups differed to the degree that their earlier SAT scores would predict.
The authors cautioned explicitly against reading the experiment as evidence that stereotype threat is the primary cause of differences in test performance between groups. That caution has been in the published record for more than twenty years, and it is still absent from most of what is written about the subject.
What the meta-analyses actually found
A great deal of research followed the original study, and enough of it now exists to be pooled and examined. The picture that emerges is of a real but modest laboratory effect whose size depends heavily on conditions, and whose literature carries the fingerprints of publication bias.
The early pooled estimate. Hannah-Hanh Nguyen and Ann Marie Ryan's 2008 meta-analysis in the Journal of Applied Psychology reported an overall mean effect size of 0.26, with genuine moderators underneath it: effects varied by whether the stereotype concerned race or gender, by how blatantly the threat was cued, and by how strongly the person identified with the domain being tested. An average with moderators that large is a signal that the conditions matter as much as the phenomenon.
The replication audit. In 2012 Gijsbert Stoet and David Geary reviewed attempts to replicate the well-known experiment on women and mathematics. Of the articles with designs that could have replicated the original result, only 55% did — and half of those were, in the authors' assessment, confounded by statistically adjusting for pre-existing mathematics scores. Among the unconfounded experiments, 30% replicated the original. Their meta-analysis found that only the studies using adjusted scores displayed the effect at all.
The publication-bias finding. Paulette Flore and Jelte Wicherts pooled 47 effect sizes from studies of schoolgirls in 2015 and found a mean effect of −0.22 that differed significantly from zero — but also “several signs for the presence of publication bias”, concluding that bias “might seriously distort the literature on the effects of stereotype threat among schoolgirls”. They called for a large replication study to produce a less biased estimate.
The replication they asked for. Three years later Flore, Joris Mulder and Wicherts ran it as a registered report — the design and analysis locked in and peer-reviewed before the data existed, which removes the room in which publication bias operates. In 2,064 Dutch high-school students aged 13 and 14 they found neither an overall effect of stereotype threat on mathematics performance nor any of the moderated effects the theory predicted.
The question that matters for real exams. Most of the literature is laboratory work, with features that would never appear in an actual testing hall. Oren Shewach, Paul Sackett and Sander Quint set out in 2019 to isolate the studies that resemble operational testing, and also identified a previously unrecognised methodological error in how studies that control for a prior test score had been analysed. Their focal sample — restricted to conditions relevant to real testing — gave an effect of d = −0.14 across 45 samples and 3,532 people. In genuine operational settings and in studies using motivational incentives, effects ran from zero to −0.14. They found nontrivial evidence of publication bias across the database, and concluded that in scenarios such as college admissions and employment testing the effect “may range from negligible to small”.
What this does not mean
A subject this contested attracts two opposite misreadings, and both are wrong.
It does not mean stereotype threat is fake or that nobody experiences it. The laboratory effect has been produced many times. Anxiety about being judged is a real thing that real people feel in real exam rooms, and a weak or uncertain average effect in a meta-analysis is not a statement about any individual's experience.
And it says nothing whatsoever about why groups differ on average on tests. This needs stating plainly, because the replication record is sometimes quoted as though it settled that question in some other direction. It does not. Evidence that one proposed explanation is weaker than claimed is not evidence for any competing explanation — it leaves the question exactly where it was. The 2004 paper that corrected the most famous overstatement about stereotype threat made precisely this point: it cautioned against treating the effect as the primary cause of group differences, which is a statement about what that experiment showed, not a claim about causes. This page takes no position on the causes of group differences in test scores, and nothing on it should be read as supporting one.
Stereotype boost and stereotype lift
Google's own results page for this topic surfaces the question “What's the opposite of stereotype threat?”, and the literature offers two answers. Stereotype boost describes performing better when a positive stereotype about your own group is made salient. Stereotype lift describes performing better because a negative stereotype is attached to some other group rather than to yours.
Both grew out of the same research tradition, were investigated with the same kinds of small laboratory experiment, and are subject to the same concerns about effect size and publication bias set out above. They are worth knowing as concepts and are not worth treating as reliable levers for changing how anyone performs.
What it means for your own test score
The practical lesson is narrower than the theory and more useful. A reasoning test measures how you performed on one occasion, under the conditions you sat it in. Tiredness, illness, a noisy room, a phone going off, an unfamiliar question format and plain anxiety all push that performance around, which is why a single number is an estimate with a margin of error rather than a fixed property of a person. That margin is discussed in how accurate IQ tests are, and what the test is actually aimed at in how IQ tests work.
It is also a good example of a pattern worth recognising in psychology generally. A striking result gets a memorable name, the name travels far beyond the evidence, and the careful version arrives years later with far less attention. The same shape appears in the research behind growth mindset, and the habits of mind that make a claim feel more certain than it is are catalogued in cognitive bias.
Why the search results and the literature disagree
There is an oddity about this topic that is visible from the search page itself, and it is the reason this article exists.
Every result on the first page for the term is a definition — teaching centres at several universities, a psychology lexicon, an encyclopedia entry, a study aid for medical-school admissions. Each explains what stereotype threat is. Google's own AI summary states it flatly as an established fact and cites, among its sources, a therapy website and a YouTube video. Of everything on that page, only the encyclopedia entry contains a section on the criticism at all.
Meanwhile Google's own “People also ask” box — on the German results page, in English — asks whether stereotype threat has been debunked, and nothing on the page answers it. That gap between what people are asking and what is published is the whole of this page's reason for being.
Trust and scope notes
This page is educational. It summarises published research on stereotype threat, including where that research disagrees with itself, and it is not psychological, medical or educational advice.
Our IQ test measures reasoning performance under the conditions in which someone takes it. It does not measure stereotype threat, it is not a diagnostic instrument of any kind, and no result from it can establish why any individual or group scored as they did.
IQ Revealed is not affiliated with, endorsed by or connected to the American Psychological Association, the Society for the Study of School Psychology, Elsevier, Taylor & Francis, the College Board, Google, or any of the journals, publishers, organisations or researchers named on this page. They are named so that each claim can be traced to the work that made it. Quoted phrases are taken from the published titles and abstracts of the sources linked above.