The reproducibility crisis: why many published scientific results are hard to replicate

A plain guide to the reproducibility crisis, its landmark replication studies, and the open-science reforms that followed.

In August 2015, a consortium of roughly 270 researchers published a paper in Science that shook the research world. Working under the Open Science Collaboration, they had tried to replicate 100 studies from three top psychology journals. The original papers had reported statistically significant effects 97 percent of the time; the replications managed that in only about 36 percent of cases. Average effect sizes in the replications were roughly half those claimed in the originals.

That project did not create the problem, but it gave the "replication crisis" its most famous number. By then, doubts had been building for years. A 2005 paper by John Ioannidis in PLOS Medicine, "Why Most Published Research Findings Are False," had argued mathematically that small samples, flexible study designs, publication bias, and financial incentives could produce literatures dominated by false positives. The psychology results turned a theoretical worry into a measurable scandal.

The issue was never confined to psychology. The Center for Open Science also ran the Reproducibility Project: Cancer Biology, trying to replicate experiments from 53 high-impact preclinical cancer papers published between 2010 and 2012. Missing reagents, undocumented protocols, and other practical barriers meant only 50 experiments from 23 of those papers could even be attempted; where replication succeeded, the effects were usually much smaller than originally reported. Similar problems surfaced in experimental economics, preclinical drug-target research, and biomedicine more broadly.

A 2016 Nature survey of over 1,500 scientists found that more than 70 percent had tried and failed to reproduce another scientist's results at least once. The rate varied by field: 87 percent of chemists, 77 percent of biologists, 69 percent of physicists and engineers, and 67 percent of medical researchers reported such failures. More than half said they had even failed to reproduce their own work.

Several structural problems keep recurring. Journals and careers reward novel, positive findings much more than null results or replication attempts. Researchers often have wide latitude in choosing samples, measures, and analyses, which can inflate false-positive rates. And many studies are published without sufficient data, code, or materials, so independent scientists cannot check them even in principle.

The response has become known as the open-science or credibility-revolution movement. Reforms include preregistration of studies, where researchers register their hypotheses and analysis plans before conducting the work, and Registered Reports, in which journals peer-review study plans before results are known. Data and code sharing have also become more common requirements. These changes aim to make science less a matter of trust and more a matter of verification.

Not everyone agrees the situation is best called a "crisis." Some scholars, such as Vazire writing in 2018, prefer "credibility revolution" to emphasize the improvements the debate has produced. Whatever the label, the 2015 psychology project, the cancer-biology replications, and the Nature survey together show that reproducibility is no longer a niche methodological concern; it is central to how modern science is judged and reformed.

People also search for

Discussion 0

Nothing has been said yet. Start it.

Log in to join the discussion

🛡️Safe SearchAlways on
Fast ResultsInstant answers
🔒Private by designYour search, your privacy