A test result arrives and it is positive. The accompanying leaflet says the test is 99 percent accurate. The obvious inference is that there is a 99 percent chance the condition is present.
That inference is not merely imprecise. Depending on how common the condition is, it can be wrong by an enormous margin, and in the case of a genuinely rare condition it can be almost exactly backwards. A positive result from a highly accurate test for something rare is more likely to be wrong than right, and no amount of improvement in the laboratory changes that, because the problem is not in the laboratory.
The reason is that accuracy figures describe how the test behaves on people whose status is already known. The question a patient has is the reverse: given this result, what is my status. Getting from one to the other requires a third number that the leaflet almost never supplies, and that number is how common the condition is in people like the one being tested.
Key takeaways
- Sensitivity and specificity describe the test; predictive values describe what a result means for a person.
- Converting one into the other requires the prevalence of the condition in the population being tested.
- When a condition is rare, even a very specific test produces mostly false positives.
- Confirmatory testing works by raising prevalence: only those who screened positive are tested again.
- Many wrong results are not statistical at all but analytical, from interference, contamination or mislabelling.
Sensitivity and Specificity Defined Cleanly
Two numbers describe how a test performs, and both are calculated by testing people whose true status has already been established by some independent reference standard.
Sensitivity is the proportion of people who genuinely have the condition that the test correctly identifies. A highly sensitive test rarely misses. Its complement is the false negative rate.
Specificity is the proportion of people who genuinely do not have the condition that the test correctly clears. A highly specific test rarely raises a false alarm. Its complement is the false positive rate.
These two properties trade against each other whenever the test produces a continuous measurement rather than a yes or no, which is nearly always. Somewhere in the range of possible values sits a cut-off, and moving it shifts the balance. Lower the cut-off and more people with the condition are caught, but so are more people without it: sensitivity rises, specificity falls. Raise it and the reverse happens. There is no setting that improves both, only settings that suit different purposes.
That is why the same measurement can be reported with different thresholds in different contexts. A test used to rule a serious, treatable condition out is set for sensitivity, accepting false alarms as the price of missing nobody. A test used to confirm before an irreversible intervention is set for specificity, accepting some misses as the price of not acting on people who are well.
The single word “accuracy” hides all of this. A test can be called 99 percent accurate because it correctly classified 99 percent of the people it was validated on, but if only one percent of that group had the condition, a test that reported negative for everyone would have scored the same.
The critical point is that neither sensitivity nor specificity depends on how common the condition is. They are properties of the test measured against a reference standard. That stability is exactly why they cannot answer a patient’s question on their own.
Why Prevalence Changes Everything

The question a person actually has is the predictive value: given a positive result, what is the probability the condition is present. Given a negative, what is the probability it is absent.
Predictive values depend on prevalence, and steeply. The mechanism is intuitive once seen. A test applied to a population produces false positives in proportion to the number of people without the condition, because those are the people who can generate one. It produces true positives in proportion to the number who have it. When almost nobody has the condition, the pool available to generate false positives is vastly larger than the pool available to generate true ones, and the false positives can outnumber the true positives even when the false positive rate is tiny.
This is not a defect in any test. It is a consequence of arithmetic, and it applies equally to laboratory tests, imaging, and screening questionnaires.
Prevalence is not a fixed property of a disease either. It is the frequency of the condition in the specific group being tested, and clinical selection changes it dramatically. A person with characteristic symptoms, a suggestive history and relevant risk factors belongs to a group in which the condition is far more common than in the general population. Testing that person is a different statistical exercise from screening an unselected crowd, using an identical test.
This is the formal justification for something clinicians do by instinct: forming an expectation before ordering the test. The pre-test probability is the prevalence that goes into the calculation, and a test result modifies it rather than replacing it. A strongly positive result in someone with no risk factors and no symptoms leaves a probability that may still be low. The same result in someone with a suggestive presentation can be close to conclusive.
A Worked Example With Real Arithmetic
Numbers make this concrete. Take a hypothetical test with a sensitivity of 99 percent and a specificity of 99 percent, which is better than most tests in routine use, and apply it to 100,000 people in three different settings.
In the first setting the condition is present in one person in a thousand, so 100 of the 100,000 have it. The test finds 99 of them. The remaining 99,900 do not have the condition, and one percent of them, meaning 999 people, test positive anyway.
| Setting | People with condition per 100,000 | True positives | False positives | Chance a positive is correct |
|---|---|---|---|---|
| Rare condition, unselected screening | 100 | 99 | 999 | About 1 in 11 |
| Moderately common, some selection | 1,000 | 990 | 990 | About 1 in 2 |
| Common in a symptomatic group | 20,000 | 19,800 | 800 | About 24 in 25 |
The first row is the one worth sitting with. A test that is right 99 times out of 100 in both directions, applied to a condition affecting one person in a thousand, produces roughly ten false positives for every true one. Someone receiving a positive result in that setting has roughly a one in eleven chance of having the condition. The test is excellent. The inference from a single positive is still weak.
The second row shows what happens when prevalence rises to one percent: true and false positives balance, and a positive result means roughly even odds. The third row shows a test used the way tests are meant to be used, on people already selected by symptoms and history, where a positive is nearly conclusive.
Negative results behave in the mirror image. In the rare-condition setting, a negative is overwhelmingly reassuring, because almost everyone testing negative genuinely does not have the condition. In the high-prevalence setting, a negative result still leaves meaningful residual probability, which is why a clinician with strong clinical suspicion may repeat or escalate testing after a negative result rather than accept it.
One more feature of the arithmetic deserves attention. Improving sensitivity barely helps the rare-condition case, because nearly all true positives were already being caught. Only better specificity reduces the flood of false positives, and each step is expensive in cost and speed.
Where Confirmatory Testing Fits
The way out of the rare-disease trap is not a better single test. It is a sequence.
Confirmatory testing works because the group entering the second test is not the original population. It is the group that screened positive, and in that group the condition is far more common than it was in the crowd. In the first row of the table, the condition affected one person in a thousand at the start; among the 1,098 people who tested positive, it affects roughly one in eleven. The prevalence has risen by nearly a hundredfold, purely through selection, and a second test applied to that group performs far better in predictive terms than the first did.
For this to work, the confirmatory test must fail differently from the screening test. If both rely on the same principle, the same interfering substance, or the same antibody, then the people who produced a false positive on the first are disproportionately likely to produce one on the second, and the second test adds much less than the arithmetic suggests. Well-designed testing pathways therefore pair methods that are mechanistically independent: an antibody-based screen confirmed by a method that separates and identifies the molecule directly, or a serological screen confirmed by nucleic acid detection.
This structure appears throughout laboratory medicine. Newborn screening programmes flag far more infants than have the conditions concerned, and the diagnostic work-up that follows is the confirmatory stage. Workplace drug testing uses an immunoassay screen followed by mass spectrometry, because the screen cross-reacts with related compounds and the confirmatory method identifies the specific molecule.
Understanding the structure changes how the first result should be received. A positive screening result is an instruction to investigate, not a diagnosis, and programmes that communicate it as one cause avoidable distress.
Analytical Versus Clinical False Positives
Not every wrong result is a statistical artefact. It is worth separating two quite different failures that both end up printed as an unexpected positive.
An analytical false positive means the measurement itself was wrong. The analyte was not present at the reported concentration, and something in the process produced a signal anyway. This is a laboratory error in the broad sense, though often not anyone’s mistake.
A clinical false positive means the measurement was correct and the interpretation was not. The substance really was elevated, but for a reason unrelated to the condition the test is used to detect. A raised cardiac troponin in someone with kidney impairment is a true measurement of a real elevation that does not indicate an acute heart attack. A positive antibody result reflecting past exposure or vaccination rather than current infection is another. So is a tumour marker raised by benign inflammation.
The distinction matters because the remedies differ. An analytical false positive is addressed by repeating the test, ideally on a fresh sample and preferably by a different method. A clinical false positive is not fixed by repeating anything: the result will come back the same, because it was right. It is addressed by reconsidering what the number means in this person.
A third category is the true positive nobody wanted. Screening can correctly identify disease that would never have caused harm in a person’s lifetime. This is overdiagnosis, a distinct problem from false positives though the two are frequently conflated.
Interference, Contamination and Sample Mix-Ups
The analytical failures have specific mechanisms, and most of them are well characterised.
Interference is the largest category. Immunoassays are especially vulnerable because they depend on antibodies binding a target, and other things in blood can bind those antibodies. Heterophile antibodies, which some people carry after exposure to animals or certain therapies, bridge the assay’s capture and detection antibodies directly, generating a signal with no analyte present. High-dose biotin interferes with a widely used assay chemistry and can push results markedly high or low depending on the assay’s design. Haemolysis, lipaemia and high bilirubin each disturb optical measurements in their own way.
Contamination is the second category, and it dominates in nucleic acid testing. Because PCR amplifies enormously, a minute quantity of DNA carried over from a previous positive sample, or from amplified product circulating in the laboratory air, produces a genuine-looking result. Laboratories control this with separated pre- and post-amplification areas, unidirectional workflow and negative controls in every batch. A contaminated run is usually caught because those negative controls turn positive.
Sample identification errors are the third and the most consequential, because the result is entirely correct for the wrong person. The controls are procedural rather than analytical: identifying the patient at the bedside, labelling at the point of draw rather than in advance, barcoding through the analytical chain, and delta checks that compare a result against that patient’s previous value. A result wildly inconsistent with everything else known about a patient should raise the identification question before the biology question.
Quality control sits underneath all of this. Laboratories run material of known concentration alongside patient samples and track the results over time, so that drift or a sudden instrument shift is caught before it corrupts a day’s work.
Questions to Ask Before Acting on One Result
The practical translation of all this is a short list of questions worth asking about any unexpected positive, and they are not statistical questions so much as contextual ones.
The first is why the test was done. A test ordered because of a specific clinical suspicion carries far more weight than one that arrived as part of a broad panel or a general health check, because the pre-test probability was higher. If nobody had a reason to suspect the condition before the result appeared, the arithmetic above applies in full force.
The second is how common the condition is in someone with this profile. Not in the population at large, but in people of this age, sex, history and symptom pattern. This single number does more to determine what a positive means than any property of the test.
The third is whether this is a screening test or a diagnostic one. Screening tests are deliberately tuned for sensitivity and are expected to generate false positives; that is the design, not a malfunction. Asking whether a defined confirmatory pathway exists, and what it involves, usually answers the question of what happens next.
The fourth is whether anything could have interfered. Recent supplements, particularly high-dose biotin, recent infections or vaccinations, medications and pregnancy are all worth mentioning, because a laboratory can often test for interference or repeat the measurement by a different method.
The fifth is whether the result fits everything else. A value that contradicts the clinical picture, the previous result and the other analytes on the same report is more likely to be an artefact than a revelation.
Frequently asked questions
If a test is 99 percent accurate, why might my positive result be wrong?
Because that figure describes how the test behaves on people whose status is already known, not what a result implies about you. The bridge between the two is how common the condition is in people like you. When it is rare, the very large group of people without it generates more false positives than the small group with it generates true ones, so most positives are false even though the test performs superbly. Nothing about the laboratory is at fault; the arithmetic is doing it.
Does repeating the test fix a false positive?
Sometimes. Repeating helps when the cause was analytical, such as contamination, a mislabelled tube or a transient interference, because a fresh sample analysed independently is unlikely to fail the same way twice. It does not help when the elevation is real but caused by something other than the target condition, since the repeat will simply confirm the same true value. It also helps less than expected if the repeat uses the identical method, which shares the same vulnerabilities.
Why do screening programmes use tests that produce so many false alarms?
Because the cost of the two errors is not symmetrical. For a serious condition that responds to early treatment, missing a case is far worse than calling someone back for further tests, so the threshold is set toward sensitivity deliberately. The programme is designed as a sequence, with a confirmatory stage that resolves the initial flags. The false alarms are a known and accepted cost of the design rather than evidence that the screening test is poor.
Can a supplement really change my blood test result?
Yes, and high-dose biotin is the well-documented example. It interferes with a widely used assay chemistry and can shift results substantially in either direction depending on how a particular assay is built, which has caused misleading thyroid and cardiac marker results. Laboratories generally advise pausing high-dose biotin before testing. Mentioning all supplements, not just prescribed medicines, gives the laboratory the chance to check for interference or use an alternative method.
What should I do if a result does not fit how I feel?
Say so, and ask whether the result is consistent with your previous values and with the other analytes on the same report. A number that contradicts everything else known about you is more likely to be an artefact than a discovery, and laboratories run automatic checks for exactly that pattern. The reasonable next steps are usually a repeat sample, a different analytical method, or a confirmatory test, rather than a decision based on the single number.
The habit these questions build is treating a result as evidence that shifts a probability rather than as an answer that establishes a fact. Most of the time the shift is decisive, because the test was ordered in the right person for the right reason and the result agrees with everything else. When it is not decisive, the appropriate response is almost always another test, a different method, or a repeat sample, rather than a decision. Very few laboratory results are urgent enough that the correct next step is to act on one number before anyone has asked whether it is likely to be true.
This is education, not medical advice. Laboratory results only carry meaning alongside your symptoms, history and examination. Talk to a qualified clinician about your own results before changing anything about your care or supplements.




