The Reproducibility Crisis Seen From the Bench

Most replication failures are not caused by fraud but by ordinary bench realities: drifting protocols, variable reagents, small samples and selective reporting.

Two identical laboratory benches side by side with matching equipment and differing printed result sheets

A postdoc spends four months trying to reproduce a published result, gets nothing, and concludes either that they are incompetent or that the original authors were dishonest. Both conclusions are usually wrong, and the fact that those are the two options that come to mind says a great deal about how poorly the mechanics of replication are taught.

The reproducibility problem became widely visible when large coordinated efforts attempted to repeat published findings in psychology, cancer biology and other fields, and a substantial fraction of the attempts did not produce the original result. Those efforts were valuable and the reaction to them was often misdirected, because the public framing settled quickly on misconduct. Fabrication exists and is damaging, but it is rare, and it explains very little of the observed pattern.

What explains most of it is ordinary. Protocols contain undocumented adjustments that the original laboratory made without noticing they mattered. Reagents differ between suppliers and between lots in ways nobody characterised. Cell lines are not the cells they are labelled as. Studies are too small to estimate an effect precisely, so the published figure is inflated by the very process that made it publishable. And the literature contains a biased sample of the experiments that were actually run.

None of these requires anyone to behave badly. Each is a predictable consequence of how bench work and publishing are organised, and each has a partial remedy that is already known.

Key takeaways

  • Failure to replicate is usually a signal about method description and effect size, not about honesty.
  • Protocols in papers are compressed summaries; the details omitted are often the ones that matter.
  • Reagent and cell line variability introduces differences between laboratories that no one recorded.
  • Small studies produce imprecise estimates, and the significant ones are systematically overstated.
  • Preregistration and registered reports address publication bias at the point where it is created.

What Replication Failure Actually Means

The word covers several distinct situations that deserve separating, because they have different causes and different implications.

Direct replication repeats the original procedure as closely as possible with new data, asking whether the same effect appears. Conceptual replication tests the same underlying claim by a different route, asking whether the idea holds rather than whether the procedure repeats. A conceptual replication that fails is ambiguous, because the new method might simply not capture the phenomenon. Computational reproducibility is narrower still: given the original data and analysis code, does the published number reappear? Failures here are pure bookkeeping problems, and they are more common than one might hope.

There is also a difference between an effect that vanishes entirely and one that appears in the same direction but much smaller. The second is far more common and is often reported as a failure, which obscures what it is telling us. A real but modest effect, first estimated in a small study that happened to catch a favourable sample, will shrink on repetition. That shrinkage is expected behaviour, not a contradiction.

Finally, a single failed replication does not refute an original finding any more than a single positive study establishes one. Both are estimates with uncertainty attached, and the replication carries its own possibility of being underpowered or of having departed from the original conditions. What settles matters is an accumulation of attempts, which is precisely what the literature has historically been poor at collecting.

Protocol Drift and Undocumented Tweaks

A stack of printed research papers with highlighted passages beside a laptop displaying a statistics plot
Illustration: Daily Lab Dish

The methods section of a paper is a compressed account written under a word limit, by people who have performed the procedure so many times that the parts they adjust automatically no longer register as decisions. This is the largest single source of replication difficulty in experimental biology, and it is almost entirely invisible from the published record.

The omitted details are rarely exotic. How long cells sit at room temperature between steps. Whether a solution is used fresh or after a week at four degrees. The exact speed and duration of a spin, and whether the brake is on. How vigorously a suspension is mixed, and whether by pipette or vortex. Which end of the incubator a plate goes in, and whether the outer wells were filled with water to reduce evaporation. Each is trivial in isolation, and collectively they can determine whether an effect appears.

Drift compounds this. Over years, a laboratory’s working version of a procedure diverges from the published one through a series of small adaptations that made sense at the time and were never written down. When a new person joins, they learn the current version by watching, which transmits the tacit adjustments but not the awareness that they are adjustments. The published protocol, meanwhile, still describes the original.

The remedy is unglamorous and effective: maintain a working protocol document under version control, record changes with dates and reasons, and publish the working version rather than a tidied reconstruction. Detailed protocol repositories exist for this. So does video, which conveys handling details that prose cannot. Where an effect turns out to depend on a step nobody documented, that dependency is itself a scientific finding worth reporting.

Reagent and Cell Line Variability

The second major source is material rather than procedural. Two laboratories following identical instructions can still be doing different experiments, because the substances they are using differ.

Antibodies are the most notorious case. An antibody is a biological product, and its performance depends on the clone, the lot, the buffer and how it has been stored and handled. Independent testing has repeatedly found commercial antibodies that do not bind the protein named on the vial, or bind it alongside several others, or work in one application and not another. A paper that cites a catalogue number without a lot number, and without evidence that the antibody was validated in that specific application, has under-specified a critical component.

Cell line misidentification is the most consequential and the best documented. Lines become cross-contaminated with faster-growing cells, most famously by a widely used cervical cancer line, and the resulting culture continues to be passaged and shared under the original name. Registers of known misidentified lines list hundreds, and papers using them continue to be published. Short tandem repeat profiling resolves this cheaply, and journals and funders increasingly require it, though adoption remains uneven.

Even correctly identified cells drift. Passage number changes gene expression, growth rate and drug response. Culture conditions, serum lot, and whether the line has been carrying an undetected mycoplasma infection all shift behaviour substantially. Serum in particular is an undefined biological mixture that varies from lot to lot.

Source of variationHow it appearsPractical mitigation
Antibody lot and specificityBands or signal differ between laboratoriesValidate in application, report lot number
Cell line identityResults diverge from the original tissue typeShort tandem repeat profiling
Passage number and culture driftEffect weakens over months in one laboratoryWorking banks, defined passage limits
Serum and media lotsGrowth or response shifts after a reorderLot testing, reserving a validated lot
Mycoplasma contaminationBroad, unexplained changes in behaviourRoutine testing of all lines
Animal microbiome and housingEffects differ between facilitiesReport supplier, diet and housing conditions

Underpowered Studies and Noisy Effects

The statistical contribution to the problem is the least intuitive and the most important, because it produces published effects that are wrong in size even when they are real in direction.

Statistical power is the probability that a study detects an effect of a given size if it exists. Power depends on the size of the effect, the variability of the measurement and the number of independent observations. Many bench experiments run with very few independent replicates, and the variability of biological measurements is considerable, so power is often low.

Low power has two consequences. The obvious one is that real effects are missed. The less obvious and more damaging one is what happens to the effects that do reach significance. In a noisy, underpowered study, only a large observed difference can clear the significance threshold, and a large observed difference arises either from a large true effect or from noise happening to fall favourably. Since noise is plentiful, published estimates from small studies are systematically inflated. The finding is then replicated at a larger sample size, the estimate shrinks toward its true value, and the replication is reported as a failure.

Two further habits amplify this. Pseudoreplication treats measurements that are not independent as though they were, counting three wells from one culture as three replicates when the culture is the unit that varies, which inflates apparent precision. And optional stopping, adding samples until the result becomes significant, converts a threshold into something that can be reached by persistence alone.

The fix is to plan sample size before the experiment, to be explicit about the unit of replication, and to report estimates with intervals rather than reducing everything to whether a threshold was crossed.

Selective Reporting and the File Drawer

Everything so far concerns individual experiments. This section concerns which experiments become visible, and it operates at the level of the literature rather than the bench.

Journals have historically preferred positive, novel, clean results. Researchers respond rationally: an experiment that produced nothing goes in a drawer, and the paper is built from the ones that worked. The published record therefore represents a filtered sample of what was done, and the filter selects precisely for the results most likely to be flukes.

Within a paper, the same selection happens at finer grain. Choices made during analysis, each defensible on its own, accumulate into a large space of possible results, and choosing among them after seeing the data inflates the chance of a false positive far beyond the nominal threshold. Which outliers to exclude, which covariates to include, which time point to feature, how to group conditions, when to stop collecting: these are all decisions with reasonable justifications available after the fact. Reporting a hypothesis formed after seeing the data as though it had been specified in advance completes the process.

The people doing this generally do not experience it as improper. Each step feels like the sensible handling of a messy dataset, and the narrative that emerges feels like the true story. That is what makes it difficult to police through appeals to integrity, and why the effective responses are structural.

Registered Reports and Preregistration

Preregistration means recording the hypothesis, design, sample size and analysis plan in a time-stamped public archive before data collection. It does not prevent exploratory analysis; it separates the exploratory from the confirmatory, so a reader can tell which claims were predicted and which were discovered in the data. Both are legitimate, and they carry very different evidential weight.

Registered reports go further and change the incentive directly. The journal reviews the introduction, methods and analysis plan before any data exist, and if it accepts, it commits to publishing the result whatever it turns out to be. Peer review therefore evaluates the question and the design rather than the appeal of the outcome, and a null result reaches print as readily as a positive one. Fields that have adopted this format report a markedly higher proportion of null findings, which is what the mechanism predicts.

The related structural remedies are worth listing together, because they attack different parts of the problem. Sharing data and analysis code allows computational reproducibility to be checked directly. Reporting checklists ensure that key methodological details, including randomisation, blinding and how sample size was determined, actually appear. Preprints and results-neutral outlets create routes for findings that would otherwise remain unpublished. Multi-laboratory collaborations run the same protocol in several places at once, which turns between-laboratory variability from an invisible confound into a measured quantity.

None of this is free. Preregistration takes planning time, sharing data takes preparation, and multi-site work takes coordination. The costs are real and much smaller than building on findings that do not hold.

Practical Steps That Improve Reproducibility

For anyone at a bench rather than designing policy, a small set of habits does most of the work, and none requires waiting for the system to change.

Write the protocol as you actually perform it, including the parts that feel too obvious to mention, and version it so changes are visible. When a procedure is modified, record what changed and why. Treat the working protocol as a living document rather than a formality to be reconstructed at writing-up.

Characterise materials rather than assuming them. Authenticate cell lines on receipt and periodically thereafter, test routinely for mycoplasma, record lot numbers for antibodies and critical reagents, and validate an antibody in the application it will be used for rather than trusting the datasheet. Keep a working bank at a defined low passage and return to it rather than culturing indefinitely.

Decide the analysis before seeing the outcome. Fix the sample size in advance based on the smallest effect that would matter, define the unit of replication honestly, state the primary outcome, and write down how outliers will be handled. Where analysis choices are made afterwards, say so, and label the result as exploratory.

Design for the confounds that operate silently. Randomise the position of samples on a plate rather than laying out conditions in blocks that align with edge effects and gradients. Blind the person measuring the outcome wherever it is feasible. Balance conditions across days and batches so that a batch effect does not masquerade as a treatment effect.

Finally, keep the record of what did not work. Failed attempts, unexpected sensitivities and conditions under which the effect disappears are the most useful information a laboratory owns, and they are almost never written down. A laboratory that knows its effect requires cells below a certain passage and a particular serum lot understands its own system far better than one that has only ever recorded successes, and it is in a far better position to help somebody else reproduce it.

Frequently asked questions

Does a failed replication mean the original result was wrong?

Not on its own. A replication is an estimate with its own uncertainty, and it can fail for reasons that have nothing to do with the original claim, including insufficient sample size, a subtle departure from the original conditions, or differences in materials that neither team knew about. What a failed replication reliably tells you is that the effect is not as robust as a single confident report suggested. The useful response is to ask which conditions differed and whether the effect depends on them, since that dependency is itself worth knowing.

Is fraud a significant part of the reproducibility problem?

It exists, it is damaging, and it appears to account for a small share of the pattern. Surveys of researchers find that admitted fabrication is rare while questionable practices such as selective reporting, dropping inconvenient data points and presenting post hoc hypotheses as predictions are common. Those practices are far more consequential in aggregate because they are widespread and often not perceived as improper. Focusing on misconduct is also strategically unhelpful, since it suggests the answer is to find bad actors rather than to change structures that produce unreliable results from honest people.

Why do some fields replicate better than others?

Largely because of effect size relative to noise, and the degree of control over conditions. A physical measurement with tight instrumental control and a large signal replicates easily. A behavioural or biological effect that is genuinely small, measured with a noisy instrument, in systems that vary between individuals and between laboratories, is much harder, and it demands larger samples and more careful specification than such studies typically receive. Field differences in publication culture matter too, particularly how readily null results appear and whether direct replications are considered publishable work.

What is the difference between reproducibility and replication?

Usage varies, but a common convention reserves reproducibility for obtaining the same result from the original data and analysis, essentially a check of the computational and bookkeeping chain, and replication for obtaining a consistent result from new data collected under the same conditions. The distinction matters because the failures have different causes and different fixes. Reproducibility failures are addressed by sharing data and code and by disciplined analysis records. Replication failures point at method description, material variability, or the size of the effect itself.

Can a small laboratory realistically do all of this?

Most of it, yes, and the highest-value items are cheap. Authenticating cell lines and testing for mycoplasma cost little relative to the wasted effort they prevent. Recording lot numbers, versioning protocols and writing down the analysis plan before running the experiment cost time only. Randomising plate layouts and blinding outcome assessment usually cost nothing at all. The expensive items are large sample sizes and multi-site work, and even there, being explicit about limited power in the write-up is more useful than presenting an imprecise estimate as though it were settled.

The most useful shift is to stop treating reproducibility as a virtue that good scientists possess and start treating it as a property that experimental systems either have or lack, for identifiable reasons. An effect that survives a change of antibody lot, a different operator and a plate laid out differently is telling you something durable. One that disappears under any of those is telling you something too, and finding out which conditions it needs is not a distraction from the science. It is the science.

Daniel Okafor Avatar