Skip to content
Resources

Sensitivity and Specificity Are Not Enough: Why Diagnostic Devices Fail Outside the Laboratory

Author
Dr. Shalini Krishnamurthy
Senior Scientist – Diagnostics
Expertise
Product Design and Development
Regulatory Strategy
Service
Medical Device Design & Development
Clinical Research (CRO)
Sector
Diagnostic Devices
Topic
Testing & Analytical Methods
Clinical & Therapeutic Devices
Published

8 min read

TL;DR

  • Sensitivity and specificity are properties of a test. Positive and negative predictive value are properties of a test in a population. A device with 95% sensitivity and 95% specificity produces mostly false positives when prevalence is 1%.
  • Pre-analytical and post-analytical errors account for approximately 70% of all errors in laboratory diagnostics — the analytical step, which is where device engineering concentrates, is the smallest contributor.
  • A one-year prospective study at an Indian tertiary care hospital covering 162,974 samples found 1,144 total laboratory errors, of which 978 (85%) were pre-analytical, dominated by clotted (0.29%) and haemolysed (0.20%) specimens.
  • Analytical validation follows a defined ladder — CLSI EP05 for precision, EP06 for linearity, EP07 for interference, EP09 for method comparison, EP17 for detection limits, EP28 for reference intervals. Skipping steps produces performance claims that do not survive review.
  • The point-of-care setting removes the trained operator, the controlled temperature, the calibration schedule, and the quality control programme. Devices designed for the lab and deployed at the point of care fail on all four.

Diagnostic devices carry an unusual burden: their output is not a therapy but a belief. A result changes what a clinician thinks is true, and everything downstream — treatment, referral, discharge, reassurance — follows from that. When the belief is wrong, the harm is caused by the actions the result triggered, not by the device itself.

That indirection makes diagnostic failure modes harder to see, and it makes the engineering discipline different from therapeutic devices in ways that are frequently underestimated.

Prevalence Determines Whether Your Excellent Test Is Useful

This is the most important and most consistently ignored fact in diagnostic device design.

Sensitivity and specificity are intrinsic properties of the test, estimated against a reference standard. They do not change with the population. Predictive value does — dramatically.

Consider a test with 95% sensitivity and 95% specificity, which most developers would consider strong performance.

In a population with 30% prevalence (a symptomatic clinic population):

  • Of 1,000 people, 300 have the condition. 285 test positive (true positives).
  • Of the 700 without, 35 test positive (false positives).
  • Positive predictive value = 285 / 320 = 89%. A positive result means something.

In a population with 1% prevalence (a general screening population):

  • Of 1,000 people, 10 have the condition. 9.5 test positive.
  • Of the 990 without, 49.5 test positive.
  • Positive predictive value = 9.5 / 59 = 16%. Five out of six positives are wrong.

Same device. Same performance. Entirely different clinical meaning.

The design consequence is that the intended use population is a design input, and it constrains the required specificity far more tightly than most teams anticipate. A device intended for screening in a low-prevalence population needs specificity in the high nineties, because every point of specificity lost multiplies false positives across the whole non-diseased population — which is nearly everyone.

This also determines the appropriate operating point. A device with a tunable threshold has a receiver operating characteristic curve, and choosing where to sit on it is a clinical decision about the relative cost of a missed case versus a false alarm. For a screening device where a false positive triggers an inexpensive confirmatory test, favour sensitivity. For a device where a false positive triggers an invasive procedure, favour specificity. This decision should be documented with clinical input, not made implicitly by whoever set the default in firmware.

The Errors Are Mostly Not Analytical

Diagnostic device engineering concentrates on the analytical step: the assay chemistry, the optics, the signal processing, the algorithm. That is where the intellectual difficulty is, and it is where the smallest share of real-world error occurs.

The laboratory medicine literature is consistent: pre-analytical and post-analytical errors account for approximately 70% of all errors in laboratory diagnostics, with pre-analytical dominating. These arise from patient preparation, sample collection, identification, transport, storage, and preparation for analysis — none of which are analytical problems.

A one-year prospective Six Sigma study at an Indian tertiary care hospital laboratory (January–December 2019) quantified this concretely: across 162,974 samples, 1,144 errors were identified, of which 978 — roughly 85% — occurred in the pre-analytical phase. The most common individual errors were clotted samples (0.29% of specimens) and haemolysed samples (0.20%). The overall process sat between four and five sigma.

The downstream consequences are measurable. Published estimates place the risk of inappropriate patient care resulting from laboratory error at 6.4% to 12%, with the probability of triggering additional unnecessary investigations substantially higher at around 19%.

Laboratory results are frequently said to influence 60–70% of clinical decisions. That specific figure is contested in the literature and its original provenance is unclear, so it is worth citing with care — but the underlying point is not in dispute: diagnostic output drives a very large share of clinical action, and errors in it propagate.

For a device developer, the actionable version is this: the specimen path is part of the device system, whether or not you designed it. A blood collection device that permits under-filling produces a citrate ratio error that no amount of analytical precision recovers. A cartridge that tolerates a 30-second delay in the lab and a 30-minute delay in a rural clinic will produce different results in the two settings. If your instructions for use say "test within 15 minutes of collection" and the deployment reality is 90 minutes, you have specified a device that cannot be used correctly.

Analytical Validation Has a Defined Ladder

CLSI evaluation protocols exist because ad-hoc performance characterisation produces claims that cannot be compared or defended. The core sequence:

  • EP05 — Precision. Repeatability (within-run) and within-laboratory precision, typically over 20 days with replicate measurement. This establishes the imprecision that every other claim inherits.
  • EP06 — Linearity across the claimed measuring interval.
  • EP07 — Interference testing. Haemolysis, icterus, lipaemia, and common concomitant medications. For point-of-care devices, this list should extend to environmental interferents specific to the deployment setting.
  • EP09 — Method comparison and bias estimation against a comparative method, analysed with Passing–Bablok or Deming regression and Bland–Altman difference plots. Ordinary least squares regression and a correlation coefficient are the wrong tools here and are routinely misused: a high r value is compatible with substantial constant and proportional bias.
  • EP17 — Limit of blank, limit of detection, limit of quantitation.
  • EP28 — Reference interval establishment or verification. Reference intervals are population-specific; transferring an interval established in one population to another requires verification, not assumption.

For qualitative and semi-quantitative tests, the relevant additions are cut-off determination, reproducibility across sites, operators and lots, and — for near-cut-off performance — the C5–C95 interval, which characterises the concentration range over which the test transitions from consistently negative to consistently positive. Devices frequently report excellent performance on high-positive and clear-negative samples while behaving unpredictably near the decision threshold, which is exactly where clinical decisions are hardest.

The Point-of-Care Setting Is a Different Device Environment

Moving a diagnostic from the laboratory to the point of care removes four things simultaneously, and each of them was doing work:

The trained operator. Laboratory technologists follow SOPs, recognise abnormal results, and know when to repeat. A ward nurse under time pressure, or a patient at home, does none of these. Human factors engineering under IEC 62366-1 is not optional for point-of-care devices; use error is the dominant failure mode.

The controlled environment. A laboratory holds 20–25 °C and moderate humidity. A primary health centre in Tamil Nadu in May does not. Reagent stability, enzyme kinetics, optical component alignment, and electrochemical sensor behaviour are all temperature-dependent. Devices must either compensate or lock out.

The calibration and QC programme. Laboratories run internal quality control daily and participate in external quality assessment schemes. A point-of-care device typically has neither. On-board controls, lot-specific calibration coding, and automatic failure detection have to substitute — and their absence is a common reason point-of-care results diverge from laboratory results in field evaluations.

The result interpretation layer. A laboratory report carries reference intervals, flags, and often a comment from a pathologist. A number on a handheld screen carries none of that. The interface has to convey uncertainty, and most do not.

Field evaluation data from India illustrates that well-designed point-of-care devices can achieve laboratory-comparable performance: comparative studies of point-of-care HbA1c devices against HPLC reference methods have reported areas under the ROC curve above 0.93, with one study reporting sensitivity of 93.6% and specificity of 97.3% against the standard assay. But the same studies also report device and test failure rates that laboratory instruments do not exhibit — a reminder that reliability in the field is a separate performance dimension from analytical accuracy.

Where Diagnostic Programmes Break Down

Validating against the wrong comparator. If the comparative method is itself imperfect, disagreement is not necessarily error. Discrepant analysis — resolving disagreements by a third method applied only to discordant samples — introduces bias and is viewed sceptically by regulators. Pre-specify the reference standard and the discordance handling before the study starts.

Reporting correlation instead of agreement. Two methods can correlate at r = 0.99 and differ by 20% at every point. Bland–Altman analysis exposes this; correlation hides it.

Reference intervals inherited rather than verified. Intervals established in a European population may not transfer to an Indian one for analytes affected by diet, body composition, altitude, or genetic factors.

Ignoring lot-to-lot variability. Reagent and consumable lot variation is a real and often dominant source of imprecision in immunoassays and lateral flow devices. Validation on a single lot understates real-world imprecision.

Software and algorithm classification underestimated. A diagnostic algorithm that produces or supports a diagnosis is regulated software. It requires an IEC 62304 lifecycle, and if it uses machine learning, it requires attention to training and test set independence, subgroup performance, and — for the FDA — potentially a Predetermined Change Control Plan if the algorithm is intended to be updated post-market.

The Framing That Helps

A diagnostic device is a measurement instrument embedded in a clinical workflow, deployed to a specific population, operated by a specific user, in a specific environment. Its performance claim is only meaningful when all four are specified.

Teams that scope the device as an analytical problem build accurate instruments that generate unreliable results. Teams that scope it as a measurement-in-context problem tend to build instruments that are slightly less elegant and considerably more useful.

RhythmRx develops diagnostic and monitoring devices for cardiac applications, with analytical validation, clinical performance evaluation, and point-of-care human factors treated as a single connected programme.

Dr. Shalini Krishnamurthy is a Senior Scientist at RhythmRx, working on analytical validation, method comparison studies and clinical performance evaluation for point-of-care diagnostic systems.

Sources

  1. Lippi G et al. — pre-analytical and post-analytical phases account for approximately 70% of laboratory errors; reviewed in Diagnostic Errors and Laboratory Medicine – Causes and Strategies, Biochemia Medica.
  2. Evaluation of pre-analytical errors using Six Sigma metrics in an Indian tertiary care hospital, January–December 2019 — 162,974 samples, 1,144 errors, 978 pre-analytical.
  3. Assessment of types and frequency of errors in diagnostic laboratories, Pathology and Laboratory Medicine International (Dove Press) — risk of improper care from laboratory error, 6.4–12%; additional inappropriate investigations, ~19%.
  4. Comparative evaluation of point-of-care and laboratory HbA1c testing in India, Cureus, 2024; Diagnostic accuracy of point-of-care HbA1c tests: a field study in India.
  5. CLSI EP05, EP06, EP07, EP09, EP17, EP28 evaluation protocols; ISO 15189:2022; IEC 62366-1:2015+A1:2020; IEC 62304.

Related resources

More from the same ground.