How Peer Review Actually Works and Where It Fails

Peer review is often treated as a certificate of truth. It is closer to an informed reading by two or three busy strangers who never see the raw data.

An editorial desk holding printed manuscript pages marked with review annotations beside a laptop showing a journal submission portal

“Peer-reviewed” is doing an enormous amount of work in public argument. It appears in press releases, in policy documents, and in the middle of online disputes as though it settles the matter. The phrase carries an implication that a group of qualified people examined the work, checked it, and found it sound.

What actually happened is narrower and stranger than that. In the typical case, an editor sent the manuscript to two or three researchers who were not paid, were working in whatever time they could find, and who read the same PDF that will eventually be published. They did not see the raw data. They did not repeat the analysis. They did not visit the laboratory. They read a description of an experiment written by the people who performed it, and they judged whether that description was plausible, complete, and interesting enough for the journal.

That is a genuinely valuable filter. It removes a great deal of confused reasoning, catches missing controls, and forces authors to acknowledge limitations they would rather leave out. It is not, and was never designed to be, a check on whether the data are real. Understanding where the line falls is the difference between citing a paper appropriately and citing it as though it were a verdict.

Key takeaways

  • Peer review assesses whether a study is plausibly designed and honestly reported, not whether its data are authentic.
  • Most reviewers see only the manuscript, not the raw data, code, or laboratory records.
  • Reviewer supply has not kept pace with submission volume, which shortens and thins reviews.
  • Different review models trade anonymity against accountability; none of them verify results.
  • Publication is the beginning of scrutiny, not the end of it.

The Path From Submission to Decision

A manuscript enters through a submission portal, where the corresponding author uploads the text, figures, supplementary material and a cover letter arguing why the work suits this particular journal. Before any scientist reads it, the submission passes through administrative checks: formatting, word limits, ethics statements, conflict-of-interest declarations, and increasingly an automated similarity scan against previously published text.

The first substantive decision belongs to an editor. At large journals this is a professional editor who may not be an active researcher; at smaller and society-run journals it is usually an academic handling submissions alongside their own work. The editor’s job at this stage is triage. A substantial fraction of submissions to selective journals never reach reviewers at all, rejected at the desk for being out of scope, too incremental, or plainly flawed. This desk rejection is fast, and it is the single largest filter in the entire system.

Manuscripts that survive triage are sent out for review. The handling editor selects candidates from the reference list, from a journal database, from suggestions the authors themselves provided, and from memory. Invitations go out. Many are declined. Some are ignored. The editor keeps inviting until enough reviewers agree, which for most journals means two, sometimes three, occasionally only one.

Reviewers then work through the manuscript and return a report plus a confidential recommendation to the editor. The recommendation is usually drawn from a fixed menu: accept, minor revision, major revision, reject, or reject with an invitation to resubmit. The editor reads the reports, weighs them, and issues a decision letter. Outright acceptance on first submission is rare. The normal path is at least one round of revision, in which authors respond point by point to each comment, and the revised manuscript returns either to the original reviewers or to the editor alone.

The whole cycle typically consumes months. Where it goes wrong is rarely dramatic. It is the accumulation of small delays: a week to find reviewers, three weeks waiting for a late report, six weeks for authors to run an extra experiment, another fortnight in the editor’s queue.

What Reviewers Are Asked to Assess

A library shelf of bound scientific journals with one volume lying open on a reading table
Illustration: Daily Lab Dish

Journals send reviewers a brief, and the brief is more specific than outsiders imagine. It asks a handful of questions that fall into three groups.

The first is about the question itself. Is it worth asking, and does the introduction show that the authors understand what is already known? A reviewer with domain knowledge is unusually good at spotting a literature review that has missed the key prior work, because they probably wrote some of it.

The second group concerns methods and internal logic. Does the design plausibly answer the stated question? Are there controls for the obvious confounders? Is the sample size defensible? Are the statistical methods appropriate to the data type, and are they described in enough detail that another laboratory could follow them? This is where competent review earns its keep. A reviewer who has run the same assay knows which failure modes the authors should have addressed and can tell when a critical control is missing from the figure legend.

The third group is about the fit between claims and evidence. Do the conclusions follow from the results presented, or has a correlation quietly become a mechanism somewhere between the results section and the abstract? Are limitations acknowledged? Is the discussion proportionate?

Alongside these, most journals ask for a judgement on novelty and suitability. That question is editorial rather than scientific, and it is the one reviewers disagree on most, because it depends on the journal’s ambitions rather than on the work.

A good review is concrete. It says which figure is unconvincing and why, which analysis should be rerun, which claim overreaches. A weak review says the paper is interesting and well written and should be published. Both take up the same slot in the editor’s inbox, and only one of them improves the work.

What Reviewers Almost Never Check

This is the section that changes how the label should be read.

Reviewers do not, as a rule, see the underlying data. They see summary tables, processed figures, and reported statistics. Unless the journal mandates data deposition and the reviewer chooses to download and inspect it, which is uncommon and unrewarded, the numbers in the manuscript are taken on trust. If a value was transcribed wrongly, the reviewer has no way of knowing.

Reviewers do not rerun analyses. Even where code is shared, executing someone else’s analysis pipeline is a day of work with no credit attached. The reported p-value, effect size and confidence interval are almost always accepted as arithmetic that happened correctly.

Reviewers do not replicate experiments. This is the most persistent misconception about the process. No journal expects a reviewer to repeat a study before publication; the cost and time would be prohibitive, and reviewers are not funded to do it. Replication happens later, elsewhere, if at all.

Reviewers do not audit laboratory records, verify that ethics approval was genuinely obtained, confirm that the reported sample size matches the number of animals or participants actually enrolled, or check that the study followed the protocol registered before it began. Some journals now compare submitted manuscripts against registered protocols, but this is an editorial check where it exists, not a reviewer task.

Finally, reviewers are poorly placed to detect deliberate fabrication. Invented data that is internally consistent and plausibly noisy reads exactly like real data on the page. Most confirmed cases of fabrication have been uncovered after publication, by readers noticing duplicated images, impossible statistical patterns, or an inability to replicate a striking result. The system that catches fraud is the scientific community reading published work over years, not the two reviewers who read it over a fortnight.

Single, Double and Open Review Models

Journals differ in who knows whose identity, and each arrangement buys something at a cost.

ModelAuthors known to reviewersReviewers known to authorsMain strengthMain weakness
Single-anonymisedYesNoReviewer can judge track record and feasibilityPrestige and affiliation bias author judgement
Double-anonymisedNoNoReduces bias from name, institution and countryIdentity often guessable from methods and citations
OpenYesYesAccountability; reviewers write more carefullyJunior reviewers hesitate to criticise senior authors
Published reportsYesOptionalReaders can see what was actually questionedAdds editorial burden and delay
Post-publication onlyYesYesFast dissemination; scrutiny continues indefinitelyNo gatekeeping before a claim enters circulation

Single-anonymised review, where reviewers know the authors but not the reverse, remains the most common arrangement in the life sciences. Its defenders argue that knowing who did the work helps a reviewer judge whether an ambitious claim is credible. Its critics point out that this is precisely the problem: the same information that makes a claim believable from a famous laboratory makes it doubtful from an unknown one.

Double-anonymised review removes author identity, which reduces certain biases in principle. In practice, specialist fields are small. Authors cite their own prior work, describe an instrument only three groups possess, or write in a recognisable style, and reviewers correctly guess identity often enough to erode the protection.

Open review, where names are attached in both directions, changes the tone of reports noticeably. Reviewers who sign tend to be more temperate and more constructive. They are also, on average, less willing to reject, and early-career researchers are understandably reluctant to write a critical report on work by someone who may sit on their next hiring panel.

Publishing the review reports alongside the accepted paper is a separate choice from revealing identities, and arguably a more valuable one. It lets a reader see which objections were raised, how the authors answered, and what the editor was persuaded by.

Reviewer Shortage and Turnaround Pressure

The system runs on unpaid labour that is expected to be delivered promptly, competently, and in addition to a full job. Submission volume has grown steadily for decades. The pool of people willing and qualified to review has not grown at the same rate, and the burden is distributed very unevenly, with a relatively small group of conscientious researchers reviewing far more than their share.

The consequences are structural rather than scandalous. Editors invite more people to secure fewer acceptances. Reviews arrive late. Some arrive short: a paragraph of impressions where a detailed methodological critique was needed. Increasingly, senior researchers delegate reviews to postdoctoral or doctoral group members, sometimes with supervision and acknowledgement, sometimes without either. A trainee’s review can be excellent, but the journal believed it was commissioning the supervisor.

Turnaround pressure comes from the other direction. Authors need publications for grants, promotion and thesis deadlines, and journals compete partly on speed. Faster decisions mean less time for reviewers to think and less appetite for demanding an extra experiment that would delay publication by months.

There is also the growth of journals whose business model depends on publishing volume. Where article-processing charges fund the operation, rejecting a manuscript costs revenue. Reputable open-access journals maintain genuine standards despite that incentive, but the incentive exists, and at the extreme end sit outlets that perform essentially no review while advertising that they do. The presence of the words “peer reviewed” on a paper says nothing until you know which journal applied them.

Preprints and Post-Publication Review

Preprint servers let authors post a manuscript publicly before, or instead of, journal review. Physics and mathematics have worked this way for decades; the life sciences adopted it more recently and more nervously.

The benefit is speed and access. A result becomes visible to everyone immediately, priority is established, and readers can evaluate it directly. During fast-moving public health situations, waiting months for review has real costs. The corresponding hazard is that a preprint carries no filter at all, and the public and press frequently cannot distinguish a preprint from a reviewed paper. A striking preprint can shape opinion for months and never survive review.

The honest framing is that a preprint and a reviewed paper differ by a specific, limited quantity of scrutiny: two or three readings by selected experts. That is worth something. It is not the difference between speculation and fact.

Post-publication review covers everything that happens after a paper appears: commentary, letters to the editor, structured critiques on public platforms, reanalysis of shared data, and the slow verdict of whether anyone can build on the result. This is where most real error correction occurs. It is also where the system is weakest institutionally. Correcting the record is laborious, unrewarded, and sometimes legally fraught. Journals are frequently slow to issue corrections or retractions, and a retracted paper continues to accumulate citations for years afterwards, often from authors who never learned it was withdrawn.

Reading a Paper With Appropriate Scepticism

None of this means published research should be discounted. It means the label should be read for what it certifies: that a manuscript was judged plausible and adequately reported by a small number of qualified readers, under time pressure, without access to the underlying data.

Practical scepticism looks like a short set of habits. Check what kind of study it is before anything else, since a case series, an observational cohort and a randomised trial support very different claims regardless of how strongly the abstract is worded. Read the methods before the discussion. Look at the sample size and ask whether the effect being claimed could plausibly be detected at that size. Notice whether the outcome reported prominently is the one the study set out to measure, or one that emerged from the analysis. Check whether the data and code are available, because groups that share them tend to have prepared them carefully.

Frequently asked questions

Does peer review mean the results have been verified?

No. It means a small number of qualified readers judged the study plausible and adequately described. Verification in the scientific sense means an independent group performing the work again and obtaining a comparable result, and that happens after publication if it happens at all. Reviewers work from the manuscript, not from the raw data, laboratory notebooks or specimens, so a result that was recorded incorrectly or reported selectively will usually pass review unnoticed as long as the write-up is internally coherent.

How many reviewers actually read a typical paper?

Two is the common number, three is generous, and one is not unusual at smaller journals or for revised manuscripts returning after a first round. The editor also reads it, sometimes closely. That means the total expert scrutiny before a paper enters the permanent literature is often a handful of hours spread across two or three people, none of whom were paid for the work or had time to interrogate it fully.

Is a paper in a prestigious journal more reliable?

On average, selective journals reject more and demand more, so quality tends to be higher. But selectivity rewards surprising results, and surprising results are more likely to be wrong, because the same statistical fluke that makes a finding striking also makes it fragile. High-profile venues have published work that later failed to replicate, and unglamorous journals routinely publish careful, durable studies that nobody publicised. Journal name is weak evidence about any individual paper.

What is the difference between a correction and a retraction?

A correction fixes an error that does not undermine the paper’s conclusions, such as a mislabelled figure, a wrong affiliation, or an arithmetic slip in a secondary analysis. A retraction removes the paper from the literature because the findings are no longer reliable, whether through honest error, undisclosed problems with the data, or misconduct. Retracted papers remain visible and marked, which matters because they continue to be cited by people who did not check.

Should preprints be ignored until they are published?

No, but they should be read with the awareness that no filter has been applied. Judge a preprint the way a reviewer would: look at the design, the sample size, the controls and whether the conclusions exceed the data. Many preprints are eventually published with only minor changes. Others never appear because review found something fatal, and the version circulating publicly is the flawed one.

Treat a single paper as one observation rather than a conclusion. The unit of scientific knowledge is not the study; it is the accumulated pattern across studies, ideally summarised by a systematic review that assessed each one’s quality explicitly. Where a striking result has not been reproduced by an independent group, the appropriate stance is interest rather than belief.

And read the paper itself where the claim matters to you. Abstracts are written to be quotable, press releases are written to be reported, and headlines are written to be clicked. The distance between what the data show and what the headline says is usually created after peer review has finished, by people the reviewers never met.

Daniel Okafor Avatar