Where Laboratory Reference Ranges Actually Come From

The printed normal range is not a boundary between health and disease. It is a statistical description of where most of a selected healthy group happened to fall.

A printed statistical distribution curve on paper beside laboratory result printouts and a calculator on a desk

Every laboratory report carries two columns: your result, and a range beside it. The range is usually labelled something like “reference interval” or, less carefully, “normal”. An asterisk or a bold H appears when your number falls outside, and that mark is what most people react to.

It is worth knowing exactly what that range is. It is not a threshold at which the body stops working properly. It is not a line agreed by a committee of physiologists after studying disease. In the overwhelming majority of cases it is a description of where the middle ninety-five percent of a group of apparently healthy volunteers fell when the laboratory, or the manufacturer of its instrument, measured them.

That construction has three consequences that follow inescapably from the arithmetic. Healthy people fall outside routinely. The range belongs to the method as much as to the biology. And a number just outside the range and a number far outside it mean quite different things, even though the report marks both the same way.

Key takeaways

  • A reference interval usually describes the central 95 percent of a healthy reference population, not a boundary of disease.
  • By construction, about one in twenty healthy results falls outside any given interval.
  • Ranges are method-specific, so the same blood can produce different numbers and different flags at two laboratories.
  • Age, sex and pregnancy shift many analytes enough that a single range would mislead.
  • Some values are governed by clinical decision limits set from outcome evidence, which are a different kind of number entirely.

The Reference Population Behind the Numbers

Building a reference interval starts with recruiting people. The formal approach asks a laboratory to define what it means by a reference individual, screen volunteers against that definition, measure them under standardised conditions, and analyse the resulting distribution. The recommended number of participants runs to at least a couple of hundred for a robust interval, partitioned further if separate ranges are needed for men and women or for different age bands.

The definition of a reference individual is where most of the judgement sits. Excluding people with the disease the test detects is obvious. Beyond that, laboratories typically exclude recent illness, pregnancy, certain medications, heavy alcohol use, extreme athletic training and sometimes smoking. Each exclusion narrows the resulting interval and makes it less representative of the population that will actually be tested, since real patients smoke, take medication and were unwell last week.

Sample collection is standardised too. Volunteers are commonly sampled in the morning after a fast, seated for a set period beforehand, with a defined tourniquet time. This matters more than it sounds. Standing up shifts fluid out of the circulation and concentrates everything measured in plasma. A prolonged tourniquet does the same locally. Several analytes vary predictably over the day. A reference interval derived from rested, fasting, seated morning volunteers does not describe the same conditions as an afternoon sample taken from someone who walked in from the car park.

In practice, very few laboratories run a full study of their own. Establishing an interval de novo is expensive, and it must be repeated whenever a method changes. The common alternatives are to adopt the interval supplied by the instrument manufacturer, to adopt one published by a professional body, or to verify a borrowed interval against a much smaller local group, typically a few dozen samples, checking only that local results do not conflict badly with the borrowed range. Verification is legitimate and pragmatic. It also means the population underlying the numbers on your report may not resemble the population your laboratory serves.

Why the Middle Ninety-Five Percent Was Chosen

Rows of archived laboratory reference manuals and binders on a shelf in a hospital laboratory
Illustration: Daily Lab Dish

Once the measurements exist, the interval is drawn around the central portion of the distribution. The convention is to take the central 95 percent, cutting 2.5 percent from each tail.

There is no biological reason for that figure. It is a compromise, and its origins are statistical convenience as much as anything. Take a wider interval, say the central 99 percent, and almost every healthy person falls inside, but the interval becomes so broad that genuinely abnormal results sit within it and go unremarked. Take a narrower one, say the central 80 percent, and real abnormalities are caught early, but one healthy person in five is flagged and the flag stops carrying information.

Ninety-five percent sits at a point where the interval is narrow enough to be informative and wide enough that most healthy people pass. It is a design choice about the trade-off between missing disease and generating alarm, made once, generically, without knowing which patient the result will belong to.

How the cut points are calculated depends on the shape of the distribution. Where measurements are approximately symmetrical and bell-shaped, the interval can be estimated from the mean and standard deviation. Many laboratory analytes are not symmetrical: they are skewed, with a long tail toward high values. Enzyme activities and several hormones behave this way. For these, the interval is either calculated after transforming the data, commonly by taking logarithms, or derived non-parametrically by simply ranking every observation and reading off the values at the 2.5th and 97.5th percentiles. The non-parametric approach makes no assumption about shape, which is why it is generally preferred, and it is also why a decent sample size matters: you cannot estimate a 2.5th percentile reliably from thirty people.

One in Twenty Healthy Results Falls Outside

This is arithmetic, not a criticism. If an interval is defined to contain 95 percent of healthy people, then 5 percent of healthy people fall outside it. Half of those are high and half are low. Nothing has gone wrong; the definition guarantees it.

The implication compounds quickly when tests are ordered in panels. Consider what happens as the number of independent analytes on a report grows.

Number of analytes reportedChance all fall inside rangeChance at least one is flagged
195 in 1005 in 100
5roughly 3 in 4roughly 1 in 4
10roughly 3 in 5roughly 2 in 5
20roughly 1 in 3roughly 2 in 3

A comprehensive metabolic panel with a blood count attached reports on the order of twenty-five values. For a completely healthy person, a report with nothing flagged at all is the less likely outcome. The table assumes the analytes vary independently, which is not quite true, since related values move together and correlation reduces the count of genuinely independent chances. Even so, the direction is unmistakable: broad panels manufacture abnormal flags in healthy people, reliably and by design.

This is the strongest argument against testing without a question in mind. Every extra analyte adds a chance of an isolated flag, and each isolated flag invites a repeat test, sometimes an imaging study, occasionally a procedure. The cascade begins with a number that was never evidence of anything.

The corollary is that how far outside matters enormously. A value a whisker beyond the cut point is the expected behaviour of a healthy population. A value several times the upper limit is a different kind of observation. The report marks both with the same asterisk, which is arguably the single most misleading piece of formatting in clinical laboratory medicine.

Method-Specific Ranges and Lab Switching

The second consequence is that intervals belong to methods. Two laboratories measuring the same analyte in the same specimen can legitimately report different numbers, because they are doing different chemistry.

Some measurements are standardised to a reference material, meaning any properly calibrated instrument should agree closely with any other. Electrolytes and glucose are broadly in this category, as is glycated haemoglobin, which has been the subject of a sustained international standardisation effort. For these, a result from one laboratory is reasonably comparable to a result from another.

Many measurements are not standardised in that way. Immunoassays, which detect a target using antibodies, are the main offenders. The antibodies in one manufacturer’s kit bind a different part of the target molecule than another’s, and where the target exists in several forms, fragments, or bound to carrier proteins, the two kits are effectively measuring different things and calling them by the same name. Thyroid hormones, several reproductive hormones, tumour markers and vitamin D assays all show meaningful between-method differences. Enzyme measurements depend on the temperature, buffer and substrate the method uses, so an activity value carries its method with it.

Because the method determines the numbers, the interval must match the method. When a laboratory changes analyser or switches kit supplier, the interval usually changes with it. Patients notice this as a result that appears to have shifted when nothing about them did.

The practical rule is that trends are only interpretable within a method. A rise across three measurements made on the same instrument at the same laboratory is a real trend. The same three numbers gathered from three different providers may be describing nothing but the differences between kits.

Age, Sex and Pregnancy Partitioning

A single interval per analyte would be wrong for a large fraction of the people tested, so intervals are partitioned where the underlying biology differs enough to matter.

Sex partitioning is routine for haemoglobin, haematocrit, creatinine, ferritin, urate and the muscle enzyme creatine kinase, among others. The differences trace to body composition, muscle mass, iron loss through menstruation and hormonal effects on production.

Age partitioning is more consequential and less well appreciated. Newborn values for bilirubin, alkaline phosphatase and many haematology parameters differ so profoundly from adult values that applying an adult range would generate nonsense. Alkaline phosphatase rises during periods of rapid bone growth, so a value that would be clearly abnormal in an adult is entirely expected in an adolescent. Lymphocytes outnumber neutrophils in young children and the relationship reverses during childhood. Kidney function estimates and several hormone levels shift across adult life as well.

Pregnancy shifts a long list of analytes, and the shifts are directional and trimester-dependent rather than random. Plasma volume expands substantially, which dilutes haemoglobin and albumin. Alkaline phosphatase rises because the placenta produces its own form of the enzyme. Thyroid binding proteins increase, which alters total hormone measurements even when the free, active fraction is unchanged. Clotting factors rise. Applying non-pregnant intervals during pregnancy produces both false alarms and false reassurance, which is why obstetric services use trimester-specific ranges where they are available.

The failure mode to watch for is a report that does not show which partition was applied. If the range printed beside a paediatric result looks like an adult range, it is worth asking.

Reference Interval Versus Clinical Decision Limit

Not every range on a report was built the way described above, and conflating the two kinds is a common source of confusion.

A reference interval is descriptive. It says where healthy people fall. A clinical decision limit is prescriptive. It says where the risk of an outcome, or the benefit of an intervention, changes enough to act on, and it is derived from studies of outcomes rather than from surveys of healthy volunteers.

Cholesterol is the clearest example. The target values quoted for cholesterol and its fractions are not the central 95 percent of the population, because the population distribution includes a great many people at elevated cardiovascular risk. They are thresholds chosen from evidence about events and treatment benefit. The diagnostic thresholds for diabetes are similar: they were set at points where the risk of specific complications rises, not at a percentile of the healthy distribution. Cardiac troponin cut-offs are defined around a percentile of a healthy population but then applied as decision limits for a specific clinical purpose, which is a hybrid arrangement.

The practical difference is what it means to be inside. Being inside a reference interval means you resemble most healthy people. Being under a decision limit means the evidence suggests your risk is acceptable. A cholesterol value comfortably within what the general population produces can still sit well above the level at which treatment reduces events.

Why Comparing Labs Directly Misleads

Putting all of this together produces a short list of habits that keep the numbers honest.

Read the result against the range printed on that report, not against a range remembered from elsewhere or found online. The range that accompanies a result is part of the result, and separating them destroys the meaning of both.

Treat a change between laboratories with suspicion until you know both used the same method. If a value looks alarmingly different from last year’s, the first question is whether the provider, instrument or kit changed. Laboratories will say if asked, and a result reported in different units is an immediate clue that something more than biology differs.

Consider the biological variation of the analyte before reacting to a small change. Some measurements are tightly regulated and vary little within a person from day to day; sodium and calcium behave this way, so even modest movement is meaningful. Others vary widely within the same healthy person across days, with several hormones, iron and some enzymes among them, and for these a difference between two samples can be entirely ordinary.

Frequently asked questions

If my result is just outside the range, should I be concerned?

Usually not by itself. A result marginally outside the interval is exactly what the design of the interval predicts will happen to a proportion of healthy people, and it is also within the range that ordinary day-to-day biological variation and analytical imprecision can produce on their own. What matters is how far outside, whether the value has moved from your own previous results, whether related analytes moved with it, and whether there is a clinical question the number is answering. A repeat sample often resolves it.

Why does my new laboratory print a different range than my old one?

Because the range describes the method, and the two laboratories are probably not using the same one. Different instrument platforms, different antibody kits and different enzyme assay conditions produce different numbers from the same specimen, and each must be paired with an interval derived or verified for that method. A change in range is therefore expected when you change provider, and it does not imply either laboratory is wrong.

Can a reference range be wrong for me personally?

It can be poorly suited to you, which amounts to the same thing in practice. Reference intervals are built from a defined group of apparently healthy volunteers, and if your physiology differs systematically from that group, through athletic training, muscle mass, ethnicity, altitude of residence, pregnancy or age, your ordinary values may sit outside the printed band without anything being wrong. This is one reason clinicians weigh your previous results more heavily than the population range.

Who decides the ranges, a regulator or the laboratory?

For most routine analytes the laboratory does, usually by adopting the manufacturer’s interval or a professional body’s recommendation and verifying it locally against a modest number of samples. Regulators and accreditation bodies require that a laboratory can justify the intervals it reports and review them periodically, but they do not generally issue a single national list of normal values. Decision limits for a handful of conditions are an exception, set by clinical guideline groups from outcome evidence.

Does a result inside the range mean I do not have the condition?

No. Reference intervals overlap between healthy people and people with early or mild disease, so a value inside the band lowers the probability of a condition without excluding it. Some diseases produce results that stay within the population range while being clearly abnormal for that individual, which is why a value that has shifted substantially from your own baseline can matter even though it never crosses a printed limit.

Finally, distinguish a flag from a finding. The flag says a number fell outside a statistically constructed band derived from other people. The finding is what that number means for one person, in context, alongside symptoms, history, other results, and previous values from the same laboratory. Those previous values are usually the most useful comparison available, because a person’s own history is a far tighter reference population than a few hundred strangers who volunteered for a study.

This is education, not medical advice. Laboratory results only carry meaning alongside your symptoms, history and examination. Talk to a qualified clinician about your own results before changing anything about your care or supplements.

Priya Raman Avatar