Quality Control Charts and the Rules That Catch Drift

A single control limit either misses real drift or stops the run every other day. Multi-rule systems exist to escape that trade-off.

A laboratory quality control station with printed control charts, vials of control material and an analyser in the background

Every analyser drifts. Reagents age, lamps dim, pumps wear, calibrators degrade, and the ambient temperature in the room is never quite what the specification assumes. None of this happens abruptly. It happens slowly enough that the instrument keeps producing plausible numbers the whole time, which is exactly what makes it dangerous.

Quality control exists to detect that drift before it reaches a patient result. The mechanism is simple in principle: run a sample of known, stable composition alongside the patient work, plot the answer, and watch for the pattern that says the measuring system has moved. The difficulty is entirely in deciding what counts as movement, because the control result varies from day to day even when nothing is wrong.

Set the decision limit too tight and the run is rejected constantly for noise, staff learn to repeat the control until it passes, and the whole system becomes theatre. Set it too loose and genuine shifts go unnoticed for weeks. Multi-rule quality control is the standard escape from that trade-off, and understanding why it is built the way it is makes the daily ritual considerably less mysterious.

Key takeaways

  • Control materials measure the whole measuring system, not the analyser alone.
  • Limits must come from your own laboratory’s data, not from the manufacturer’s insert.
  • A single wide limit misses real error; a single narrow limit rejects good runs constantly.
  • Multi-rule systems combine a sensitive trigger rule with confirmatory rules to keep false rejection low.
  • The pattern of the violation, not the fact of it, points to the cause.

What Control Materials Represent

A control material is a sample of stable composition, usually a lyophilised or frozen serum-like matrix with analytes present at a known approximate concentration, run exactly as a patient sample would be. It goes through the same pipetting, the same reagents, the same optics and the same calculation. That is the point: it interrogates the entire measuring system rather than any single component.

The distinction between a control and a calibrator matters and is often blurred. A calibrator sets the relationship between the instrument’s raw signal and the reported concentration. A control checks whether that relationship still holds. Using the same material for both destroys the check, because the system will always agree with the material it was set by. Controls should be independent of calibration wherever possible, ideally from a different manufacturer than the reagent, so that a reagent lot problem shows up rather than being absorbed silently.

Concentration selection carries real weight. Controls are run at two or three levels chosen to sit near medically important decision points, not at convenient round numbers. A glucose control at a value nobody makes decisions about says far less than one near the threshold used to diagnose diabetes. Imprecision also varies with concentration, so a system performing acceptably mid-range may perform poorly at the low end where detection limits bite.

Commutability is the subtler issue. A control behaves like a patient sample only up to a point. Manufacturers stabilise these materials, and the additives, the protein matrix and the lyophilisation itself can make them respond differently to a method change than real serum would. A control can therefore look stable through a reagent lot change that shifts patient results, or shift when patient results have not moved. This is why patient-based checks, such as monitoring the average of patient results over time, complement rather than duplicate control charting.

Building a Levey-Jennings Chart

A monitor showing a control chart with daily plotted points and marked standard deviation limits
Illustration: Daily Lab Dish

The chart is a plot of control results over time, with the horizontal axis being run or day and the vertical axis the measured value, marked with a central line at the established mean and horizontal lines at fixed multiples of the standard deviation.

Establishing those lines is the part most often done badly. The mean and standard deviation must come from the laboratory’s own data, generated on its own instrument, by its own staff, using the control lot actually in use. Values printed on the manufacturer’s insert are derived across many instruments and sites, so they almost always span a wider range than a single stable analyser produces, and adopting them as chart limits produces a system that never flags anything.

A working estimate usually needs data from at least twenty separate runs on separate days, because same-day replicates capture only within-run variation and miss the day-to-day component that dominates real performance. Recalculating after a longer period, once several months of data exist, gives a more honest standard deviation.

When a new control lot arrives the process partly repeats, because lots differ and the target value shifts. Good practice is to run the new lot alongside the old, establish the new mean from that parallel data, and carry across the standard deviation from the established lot rather than recalculating it from a short and unrepresentative run.

Reading the chart is a matter of looking for structure rather than at individual points. A well-behaved chart shows points scattered randomly on both sides of the mean, mostly within one standard deviation, with no runs, no slopes and no clustering. Anything that looks organised is a signal, because random variation does not produce organisation.

Why One Rule Is Never Enough

Suppose the only rule is that a control result outside three standard deviations rejects the run. On a roughly normal distribution a stable system produces such a point very rarely, so false rejections are rare. The problem is what the rule misses: a systematic shift of one and a half standard deviations moves the whole distribution while still leaving most points inside the band. Detection is occasional, and a laboratory could run for weeks with a real bias present.

Tighten the rule to two standard deviations and detection improves substantially, but roughly one in twenty results from a perfectly stable system falls outside that limit by chance. With two control levels run daily, that means a false rejection every few days. The predictable human response is to repeat the control, get a pass, and carry on, which quietly converts the whole system into a formality.

Rule used aloneSensitivity to a real shiftFalse rejection burdenPractical outcome
Outside 3 SDLow for shifts under 2 SDVery lowReal bias persists undetected
Outside 2 SDModerateHigh with multiple controlsRepeat-until-pass culture
Outside 2 SD, twice consecutivelyGood for systematic shiftLowSlower to trigger, but reliable
Range between two levels exceeds 4 SDGood for random errorLowDetects imprecision, not bias

The table points at the resolution. No single rule is simultaneously sensitive and quiet, but a combination can be. Use the sensitive limit as a warning that triggers examination rather than rejection, then apply further rules that a stable system is very unlikely to violate. The two-standard-deviation exceedance stops being a verdict and becomes a question.

Common Multi-Rule Combinations

The widely used multi-rule framework is conventionally written in a compact notation: a number of control observations, then the limit they must exceed. Reading the notation is most of the battle.

A warning rule flags a single observation beyond two standard deviations. On its own it rejects nothing; it directs attention to the remaining rules.

The first rejection rule flags a single observation beyond three standard deviations, indicating large random error or a substantial shift, and it is almost never a chance event on a stable system.

The second flags two consecutive observations beyond the same two-standard-deviation limit on the same side. Two independent chance excursions in the same direction are unlikely, while a systematic shift produces them readily, which makes this the workhorse rule for detecting bias.

The third flags a difference of four standard deviations between two control observations within a run, typically one above the mean and one below. It is a range rule and responds to imprecision rather than bias, since a systematic shift moves both levels together and leaves the gap unchanged.

The fourth flags four consecutive observations beyond one standard deviation on the same side, and the fifth flags ten consecutive observations on the same side of the mean at any distance. Both are slow and insensitive to sudden failure, and both excel at catching gradual drift that never produces a dramatic point.

Not every laboratory needs the full set. Where a method is very precise relative to the clinical requirement, a wide margin exists between analytical performance and what would actually mislead a clinician, and a simpler set of rules is adequate. Where a method is marginal for its purpose, more rules and more controls are needed to achieve acceptable error detection. Choosing rules to match method capability, rather than applying the same set everywhere, is the mark of a mature quality programme.

The shape of a control chart violation is diagnostic, and reading it saves a great deal of blind troubleshooting.

A shift is an abrupt change in level, where the mean moves to a new value and stays there. The chart shows a step. Because the change is abrupt, the cause is almost always something discrete that happened at that moment: a new reagent lot, a recalibration, a new control lot, a maintenance intervention, a replaced part, a software update. Establishing what was done just before the step usually identifies the cause faster than any amount of instrument investigation. This is precisely why maintenance and lot change logs are valuable, and why undocumented interventions are so corrosive.

A trend is a gradual movement in one direction over many runs. Because the change is progressive, the cause is almost always something that degrades continuously: a lamp losing output, an electrode ageing, a reagent slowly deteriorating, a control material evaporating or degrading in storage, a pump losing volume accuracy as tubing fatigues, or a temperature control loop slipping. Trends are the failure mode that single-point rules handle worst and that consecutive-observation rules handle best.

Increased scatter without any change in mean is a third pattern, and it points to random rather than systematic error: bubbles in a reagent line, a partially blocked probe, unstable temperature, imprecise pipetting, or inconsistent sample handling. The mean stays put while the points spread out.

A fourth pattern is frequently misread: control values look fine while clinicians question patient results. That points to a commutability problem, and investigating it needs patient-based evidence such as split samples with another laboratory or a review of population result averages.

Responding to an Out of Control Run

The instinct on a rejection is to repeat the control immediately. Occasionally that is right, but as a default it is the single most damaging habit in laboratory quality management, because a repeat that passes provides no information about whether the first result was real.

A more defensible sequence starts with containment. Results generated since the last acceptable control run are held rather than released, because that is the window in which the error could have affected patients. Deciding the extent of the window is a judgement, and it is much easier when controls are run frequently.

Next comes examination of the rule violated and the pattern involved. A range violation and a consecutive-shift violation point in different directions, and the investigation should follow the pattern rather than proceeding through a generic checklist.

Then come the cheap checks: reagent condition and expiry, control handling and reconstitution, instrument alarms, the maintenance log, and whether anything changed since the last good run. Many rejections resolve here, most often through a control vial reconstituted inaccurately or left out too long.

Only then does repeating make sense, and it should use a fresh vial rather than the same one, so a control handling problem is distinguished from an instrument problem. If the fresh vial passes and control handling explains the failure, the run can be released with that documented. If it fails, the problem lies in the measuring system, and corrective action follows: recalibration, part replacement, reagent lot change or engineer intervention.

Last comes the disposition of held results. Once the system is demonstrably back in control, results in the affected window are re-run where sample stability permits, and anything already released that would change clinical interpretation must be corrected and communicated. Documenting the sequence is not bureaucracy; it is the evidence that the laboratory noticed, contained and fixed a real problem.

Reviewing Charts for Slow Deterioration

Daily quality control asks whether today’s run is acceptable. It is poorly suited to the separate and equally important question of whether the method is as good this quarter as it was last year, and answering that requires stepping back from individual points.

The practical tool is periodic review of summary statistics. Calculating the monthly mean and standard deviation for each analyte and each control level, and plotting those summaries over a year, exposes movement that daily charting hides. A standard deviation creeping upwards month by month says precision is deteriorating even while every individual run passes. A monthly mean drifting steadily away from the established target says bias is accumulating.

Comparing the laboratory’s own long-term mean against an external comparison programme adds the dimension that internal control cannot supply. Internal control measures consistency against the laboratory’s own history, and a system that has drifted slowly and consistently will look entirely stable against its own drifted baseline. External quality assessment, where the same material is analysed by many laboratories and results compared, is the check on that blind spot.

Rejection frequency is a third dimension. A method rejecting far more often than expected is either unstable or has limits set too tightly. A method that has never triggered a rule in a year more likely has limits set too widely than performs flawlessly.

The reason all this matters is that limits set once and never revisited slowly stop describing the system they were derived from. Recalculating from accumulated data, updating when a method genuinely improves, and resisting the temptation to widen limits after an inconvenient rejection are what keep the chart connected to reality.

Frequently asked questions

Can the manufacturer’s stated range be used as chart limits?

It should not be, other than as a temporary starting point for a brand-new method with no local data. Insert values are derived across many instruments, sites and operators, so they describe the spread of a whole user base rather than the performance of one stable analyser. Limits built on them are typically far too wide, producing a chart that flags almost nothing and a quality system that gives false reassurance. Collect local data over at least twenty separate days, calculate your own mean and standard deviation, and treat the insert range purely as a check that the local mean is plausible.

How many control levels and how often should they be run?

Frequency is driven by how much patient work would need repeating if a failure were detected, and by the stability of the method. Running controls more often narrows the window of results at risk, which is the main practical benefit. Two levels covering the clinically relevant range is the common minimum, with a third where decisions span a wide range of concentrations. Accreditation requirements set a floor, but the floor is rarely the right answer on its own, and less stable methods justify more frequent checks.

Is it ever acceptable to repeat a control and move on?

Only when a specific, documented cause explains the first result and the repeat confirms that the system itself is sound. A vial reconstituted with the wrong volume, a vial past its open-container stability, or an obvious pipetting error are genuine explanations. Repeating simply because the number was inconvenient, with no cause identified, is not. The distinction is whether an explanation was found before the repeat or invented after it, and the honest test is whether the reasoning would survive being written down in the log.

What happens when a new reagent lot shifts the control values?

A step change at a reagent lot change is common and does not automatically mean either lot is faulty. Good practice is to run the new lot in parallel with the old before switching, so the size of any shift is known in advance rather than discovered afterwards. If the shift is small relative to the method’s allowable error, the new lot can be adopted with the chart target adjusted and the change documented. If the shift is large enough to alter clinical interpretation, the lot needs investigation with the manufacturer and patient-based comparison before use.

Do the same principles apply to point-of-care devices?

The principles do, though implementation differs. Point-of-care instruments are often operated by staff without laboratory training, in settings where a rejected run has immediate clinical consequences, and many use single-use cartridges with built-in checks rather than liquid controls. Charting still applies wherever liquid controls are run, and the logic about local limits and pattern reading holds. The added risk is that devices are distributed and lightly supervised, so a central review across all devices is often more informative than any single device’s chart.

What separates a functioning quality system from a ritual one is looking at the chart rather than the pass or fail flag. A run can satisfy every rule while the last fifteen points sit quietly above the mean, and that pattern is telling you something the rules have not caught up with. Spend a minute a day on the shape of the plot, keep a legible log of every intervention so steps can be matched to causes, and review the summary statistics quarterly. Those three habits catch most of what rule checking alone misses.

Tom Bradbury Avatar