The Journal24 August 20267 min read
Five famous studies, five different endings
Where the repetition failed — and where it succeeded
Daryl Bem needed no unusual means to get his result. In 2011 the Journal of Personality and Social Psychology, one of the field’s most respected journals, printed his paper concluding that people can anticipate future events. The statistics were standard, the samples were the usual size, the reviewers found no formal error — and the result was scientifically highly implausible. That left a question in the room more uncomfortable than any single fraud: if the usual methods can demonstrate precognition, what else do they demonstrate? Three independent attempts at repetition by Ritchie, Wiseman and French found none of it again in 2012.
The answer came four years later. The Open Science Collaboration, a network of more than two hundred researchers, took on the 2008 volumes of three leading journals, a pool of 488 articles. From it they carried a hundred repetitions through to the end, as close to the originals as possible, often using the first authors' materials. 97 % of the original findings had been significant. Roughly a third of the originally significant findings could be shown again, 36 % to be exact, and the effect sizes fell on average from r = .403 to r = .197 — to about half. Since then, every textbook chapter has had the question hanging over it: which half is it in?
The causes of that number are known and unspectacular: undersized samples, analyses with many open decisions, a publication practice that prints only significant results. How much the second cause achieves on its own was demonstrated by Joseph Simmons, Leif Nelson and Uri Simonsohn in 2011. Whoever decides freely during the analysis when to stop collecting and which control variables and conditions to report can make almost any hypothesis look significant. None of these causes is fraud; together they produce a literature in which too many findings look too good. The remedies are equally unspectacular, and they work: preregistration takes the open decisions out of the analysis, shared data makes recalculation possible, many-lab studies replace the single laboratory with a network. Many Labs 2 showed in 2018 what that looks like — twenty-eight classic findings, 186 researchers, 60 laboratories, 125 samples and a clean separation between what holds and what does not.
Two misunderstandings deserve sentences of their own. The first: that the replication crisis proves people lie in psychology. Exposed frauds did in fact provide the overture — the Stapel case is the best known. But the tests of 2015 and 2018 are not about deceit; they are about incentives. A field that rewards surprise and not repetition gets surprises. The second: that the crisis is a peculiarity of psychology. It is not, though the rates come out unevenly. In preclinical cancer research an Amgen team confirmed only 6 of 53 landmark studies it re-examined; in experimental economics 11 of 18 held — markedly worse and markedly better than in psychology. Five famous cases from this archive show how differently the endings turn out.
| Case | How it was tested | Outcome |
|---|---|---|
| Open Science Collaboration, 2015 | 100 replications from three journals | 36 % significant again, effects from r = .403 to r = .197 |
| Ebbinghaus, forgetting curve | 2015, by the original recipe, one participant | passed — only the 31-day interval diverged |
| Kahneman, framing effect | Many Labs 2, dozens of countries | held |
| Kahneman, priming chapter | no replication needed | withdrawn by the author himself |
| Stanford prison | archival analysis instead of a laboratory | fallen — instructions to the guards |
| Festinger, one dollar against twenty | hundreds of variants | the core held, the most elegant choice paradigms did not |
| Milgram, obedience | Burger 2009, up to the 150-volt mark | 70 against 82.5 per cent — no meaningful difference |
Ebbinghaus’s forgetting curve: passed
The archive’s oldest finding is also its most literally tested. In the 1880s Hermann Ebbinghaus learned meaningless syllables and measured his own forgetting. In 2015 Jaap Murre and Joeri Dros repeated the experiment to the original recipe — a single subject, Dros himself, then twenty-two years old, roughly 70 hours of data collection — and laid their curve over the one from 1885. The shapes matched; only at the 31-day interval did the new measurement diverge, which the authors openly record as unexplained. A study with a sample size of one passed its repetition after a hundred and thirty years, because its procedure lay fully open. Replicability, that is the first lesson, is not a question of case numbers alone but of transparency.
Kahneman: one finding held, one chapter did not
Kahneman and Tversky’s prospect theory is among the crisis’s winners: Many Labs 2 tested the framing effect across dozens of countries, and it held. At the same time a chapter fell out of Kahneman’s own bestseller — the priming findings, to which he had given generous space, proved undersized and poorly replicable. Kahneman drew the conclusion in two steps. In 2012 he wrote an open letter to priming researchers demanding independent repetitions, and it began with a warning:
I see a train wreck looming. … your field is now the poster child for doubts about the integrity of psychological research
In 2017 a public statement followed. There he conceded that he had put too much faith in the undersized studies; the effects could not be as large and as robust as his chapter had presented them. The case shows the second lesson: the crisis does not run between people but between findings. The same author can stand on both sides.
The Stanford prison: fallen, but not in the laboratory
The Stanford prison experiment was never a study in the strict sense — no control group, the investigator acting as prison superintendent in the middle of the events. What brought it down was not a laboratory but an archive. Thibault Le Texier worked through the holdings at Stanford — in 2018 as a book, in 2019 in the American Psychologist — and found briefings given to the guards that the textbook version of spontaneous role decay did not survive. The BBC study, broadcast in 2002 and published in 2006, had already set the arrangement up afresh under independent supervision, explicitly not as a replication but as a counter-test — and had delivered a different result. Third lesson: some findings fall not in the laboratory but in the filing cabinet.
Festinger’s dissonance: the core held, the edges did not
Festinger’s experiment of 1959 — one dollar against twenty — counts as robust: the basic pattern, that the smaller justification forces the larger change of attitude, has been recovered in hundreds of variants. But the crisis sorted things here too. The free-choice paradigms that were meant to show cognitive dissonance with particular elegance came under suspicion, because part of their results can rest on a statistical artefact of the design. The core stock remained; the family shrank. Fourth lesson: even confirmed theories lose outlying districts under testing.
Milgram’s obedience: replicated and reappraised all the same
Milgram’s obedience experiments were repeated by Jerry Burger in 2009 in a softened form — for ethical reasons, only up to the 150-volt mark. In Burger’s baseline condition 70 % of participants went that far, against 82.5 % in Milgram; the difference is not significant. At the same time Gina Perry’s archival work showed how far the experimenter departed from the script and how incomplete the debriefing was. What remains is narrower and harder at once: the finding exists, its reading as proof of blind obedience does not. Fifth lesson: replication and reappraisal are two different tests, and a study can pass one and lose the other.
What remains
Secure is the number from 2015 and its diagnosis. A third of the tested findings held, the effects shrank by roughly half, and the causes lie in the working practice of the field, not in the character of its people. What remains disputed is how far the number generalises. It comes from three journals of a single year, it says nothing about the textbooks as a whole, and still less about findings that were never tested — which is most of them.
The five cases also show that “holds” and “falls” are too coarse. A finding can be replicated and reinterpreted at once. A core can hold while its most elegant variants break away. And a study can fail in the archive without ever having been tested in the laboratory. In everyday life, findings that carry weight can be recognised by marks that are learnable: large or numerous samples, preregistered analysis, independent repetitions, effects that do not shrink with every new study. Add to that authors who revise their own numbers downward when the data demand it. Every file in this archive carries a replication status. It is not a grade for the person but a statement about the evidence, and it changes when the evidence changes. That is exactly what separates a crisis from a collapse: a field that re-tests its own findings and publishes the results is working.
Sources, and why they are here
Bem, D. J. (2011). Feeling the future: Experimental evidence for anomalous retroactive influences on cognition and affect. Journal of Personality and Social Psychology, 100(3), 407–425.
The trigger — not because of an error but because no formal error could be found. That is exactly what made the question uncomfortable.
Ritchie, S. J., Wiseman, R., & French, C. C. (2012). Failing the future: Three unsuccessful attempts to replicate Bem's „retroactive facilitation of recall“ effect. PLoS ONE, 7(3), e33423.
Three independent repetitions, three times nothing. And a lesson in passing: the journal that printed Bem rejected this paper.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.
Puts a number on what the second cause achieves on its own: anyone deciding freely during the analysis can make almost any hypothesis look significant.
Open Science Collaboration (2015). Estimating the reproducibility of psychological science. Science, 349(6251), aac4716.
The figure at issue — and its limits: three journals, one year, a hundred replications.
Klein, R. A., Vianello, M., Hasselman, F., et al. (2018). Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science, 1(4), 443–490.
The answer in the form of a procedure: 28 findings, 60 laboratories, 125 samples — and a clean separation of what holds from what does not.
Murre, J. M. J., & Dros, J. (2015). Replication and analysis of Ebbinghaus' forgetting curve. PLOS ONE, 10(7), e0120644.
The first of the five cases: a study with a sample size of one that passed after a hundred and thirty years — because its procedure lay fully open.
Kahneman, D. (26 September 2012). Open letter to priming researchers; and his statement on Schimmack's priming analysis (February 2017).
The source of the quotation — and the second of the five cases: the same author stands on both sides of the crisis, with one finding that held and one chapter withdrawn.
Le Texier, T. (2019). Debunking the Stanford Prison Experiment. American Psychologist, 74(7), 823–839.
The third case, and the only one that failed not in the laboratory but in the filing cabinet: instructions to the guards that the textbook version did not survive.
Reicher, S., & Haslam, S. A. (2006). Rethinking the psychology of tyranny: The BBC prison study. British Journal of Social Psychology, 45(1), 1–40.
The counter-test under independent supervision — expressly not a replication, and with a different result.
Burger, J. M. (2009). Replicating Milgram: Would people still obey today? American Psychologist, 64(1), 1–11.
The fifth case: 70 against 82.5 per cent up to the 150-volt mark, the difference not meaningful. The finding exists — its reading as blind obedience does not.
Begley, C. G., & Ellis, L. M. (2012). Raise standards for preclinical cancer research. Nature, 483(7391), 531–533.
The comparison outward and downward: 6 of 53 re-examined preclinical studies confirmed. The crisis is no peculiarity of psychology.
Camerer, C. F., Dreber, A., Forsell, E., et al. (2016). Evaluating replicability of laboratory experiments in economics. Science, 351(6280), 1433–1436.
The comparison upward: 11 of 18. Together with the cancer preclinical work it shows that the rate depends on the field and not on science.