AK Interactive Documentary · How to Read Clinical Journals
Episode II · CMAJ 1981 · Vol. 124 · pp. 703–710

To Learn the Features of
A Diagnostic Test

Sensitivity · Specificity · Likelihood Ratios · Spectrum Bias
An interactive exploration by Alexandra Kitty · alexandrakitty.com
AK Perspective Historical Context The Three Questions The 2×2 Table Probability Calculator The Biases Quiz True Crime Bridge

The Test That Isn't Quite a Test

Part II arrived about a week after Part I in our reading pile at McMaster, and I remember thinking the title was almost deceptively modest. "To learn the features of a diagnostic test." As if a test either works or it doesn't, and all we had to do was learn which was which.

What Sackett actually delivered was a controlled demolition of medical certainty. By the time I finished it, I understood that nearly every clinical test — including many that physicians ordered daily without a second thought — was operating in a probabilistic fog that most practitioners never acknowledged. The 2×2 table felt like a little bomb. How could something so simple expose so much?

The piece that stayed with me longest was the discussion of spectrum bias — the way a test validated on dramatic cases (obvious disease vs. obviously healthy) could completely collapse in the messy middle where it actually needed to work. I kept thinking about it years later when I started covering true crime cases. The "test" in those cases was often eyewitness testimony, or bite-mark analysis, or hair microscopy. And the "validation cohort" was almost never the population the test would face in a real investigation. They'd perfected the tool on perfect conditions, then deployed it in chaos.

The box in Hamilton is still packed. But I don't need to unpack it to reconstruct the argument — I've been applying it for forty-five years.

Part II of the How to Read Clinical Journals series appeared in the Canadian Medical Association Journal on March 15, 1981, in Volume 124, pages 703–710. Like Part I, it emerged from the McMaster University clinical epidemiology group that David Sackett was building into one of the most influential methodological research programs of the twentieth century. The article asked a deceptively simple question: when you read a paper about a diagnostic test, how do you decide whether the test is actually any good?

The answer, as with the rest of the series, was organized around a structured set of critical questions — questions that, once internalized, permanently change how you read any claim that a test "works."


The Problem That Drove the Paper

1978 · The Paper That Lit the Fuse

Ransohoff & Feinstein: "Problems of Spectrum and Bias in Evaluating Diagnostic Tests"

Three years before Part II appeared, Alvan Feinstein and David Ransohoff published a landmark critique in the New England Journal of Medicine identifying two structural problems that undermined most published diagnostic test studies: spectrum bias and workup bias. Sackett built Part II directly on their framework, transforming it into a teachable set of appraisal criteria that a medical student — or a journalist, or a juror — could actually apply.

By 1981, diagnostic testing was proliferating. Biochemical analyzers were making it possible to run panels of 20 or 30 tests simultaneously on a single blood draw. Clinicians were ordering tests in industrial quantities, often without a clear sense of what a positive result actually meant in a patient with a moderate — not dramatic — probability of disease.

The central problem was Bayesian, though it was rarely framed that way. A test's ability to correctly identify disease (sensitivity) and correctly exclude disease (specificity) are properties of the test itself. But what a clinician actually needs to know — the probability that their specific patient has the disease, given a positive test result — depends critically on how common the disease is in that patient's context. The same test could be highly informative at a specialist referral centre and nearly useless in a primary care screening setting.

"A test with 95% sensitivity and 95% specificity will produce 50 false positives for every true positive when the disease prevalence is 1 in 1,000."

This counterintuitive arithmetic was not widely understood in 1981 — and, as COVID-19 rapid antigen testing controversies in 2020–2021 demonstrated, it remains poorly understood today. When COVID antigen tests showed 98–99% specificity but only 41% sensitivity in asymptomatic people, the positive predictive value in low-prevalence settings dropped to around 33%, meaning two-thirds of positive results in asymptomatic screening programs were false positives.

Sackett's Part II gave readers the tools to see this coming — four decades before the pandemic made it front-page news.

The Legacy Thread · 1981 → Today

From CMAJ to STARD, Cochrane, and the Fagan Nomogram

The structured appraisal framework in Part II directly seeded the QUADAS (Quality Assessment of Diagnostic Accuracy Studies) tool developed in 2003, and the STARD (Standards for Reporting of Diagnostic Accuracy) guidelines — now mandatory at leading medical journals. The "Users' Guide to the Medical Literature" chapter on diagnostic tests, which Sackett's McMaster group produced through JAMA in the 1990s, translated these principles to a new generation. Today, every systematic review of a diagnostic test conducted by the Cochrane Collaboration uses a checklist that traces its lineage back to the three questions in this 1981 paper.


The Three Questions for a Diagnostic Test Study

Sackett organized the appraisal of a diagnostic test study around the same three master questions from Part I — validity, results, and applicability — but populated each with criteria specific to the diagnostic setting.

Question One · Validity
Are the results valid?
  • ⬡ Was there an independent, blind comparison with a reference (gold) standard?
  • ⬡ Was the test evaluated in an appropriate spectrum of patients — including mild and ambiguous cases, not just obvious ones?
  • ⬡ Did all patients receive both the index test and the reference standard regardless of result (avoiding verification bias)?
  • ⬡ Was the reference standard applied without knowledge of the index test result (and vice versa)?
Question Two · Results
What are the results?
  • ⬡ What are the sensitivity and specificity of the test?
  • ⬡ Are likelihood ratios presented (or calculable), allowing probability revision?
  • ⬡ How precise are the estimates — are confidence intervals reported?
  • ⬡ Is the test result reproducible across observers and settings?
Question Three · Applicability
Will the results help me care for my patients?
  • ⬡ Is the test available, affordable, and accurate in my setting?
  • ⬡ Can I estimate the pre-test probability for my patients from local data?
  • ⬡ Will the result meaningfully change management — is there a treatment threshold to cross?
  • ⬡ Will patients receive the treatment if the test is positive — is it actually available?

The 2×2 Contingency Table

Every diagnostic test study reduces to a 2×2 contingency table that cross-classifies test results against true disease status. From four cells — true positives (a), false positives (b), false negatives (c), and true negatives (d) — every operating characteristic of the test can be derived.

Disease Present Disease Absent
Test Positive a
True Positive
Test says YES,
disease IS there
b
False Positive
Test says YES,
disease is NOT there
Test Negative c
False Negative
Test says NO,
disease IS there
d
True Negative
Test says NO,
disease is NOT there

Sensitivity

The test's ability to detect disease when disease is present. A highly sensitive test rarely misses cases.

Sensitivity = a / (a + c)
→ Rules OUT disease when negative (SnNOUT)

Specificity

The test's ability to exclude disease when disease is absent. A highly specific test rarely raises false alarms.

Specificity = d / (b + d)
→ Rules IN disease when positive (SpPIN)

Positive Predictive Value

Of everyone who tests positive, the proportion who actually have the disease. Critically dependent on disease prevalence.

PPV = a / (a + b)

Likelihood Ratio (+)

How much more likely a positive test result is in a person with disease vs. without. The most portable measure — not affected by prevalence.

LR+ = Sensitivity / (1 − Specificity)
SnNOUT

High Snsitivity + Negative result = rules OUT disease. Use highly sensitive tests to screen — a negative result is reassuring.

SpPIN

High Specificity + Positive result = rules IN disease. Use highly specific tests to confirm — a positive result is meaningful.


Post-Test Probability Calculator

This is the heart of what Part II was teaching: a positive test result means very different things depending on the pre-test probability (how likely was disease before the test?). Use the sliders to explore how pre-test probability and likelihood ratio combine to determine your patient's actual post-test probability of disease.

⚙ Fagan Nomogram Simulator

30%
5.0
Post-Test Probability
68%
Moderate-to-high probability. Likely warrants treatment or further confirmatory testing.
Pre-Test Odds
0.43
Post-Test Odds
2.14
Probability Shift
+38%

Notice the asymmetry: when pre-test probability is low (say 5%), even a very high LR+ of 10 only brings post-test probability to ~35%. This is the base rate problem that drove massive COVID testing controversy in 2020–2022. When COVID antigen test sensitivity in asymptomatic persons was 41% and specificity 98.4%, the positive predictive value in a low-prevalence setting fell to just 33% — meaning two-thirds of positive screening results were false positives.


The Biases That Corrupt a Diagnostic Test Study

Sackett's 1979 taxonomy of biases in analytic research, and the 1978 Ransohoff–Feinstein paper on spectrum and bias in diagnostic tests, formed the backbone of the validity checklist in Part II. Each bias below inflates the apparent performance of a test beyond what it will deliver in real-world clinical use.


Why Spectrum Bias Matters: A Diagram

The same test can appear near-perfect in a research study and near-useless in the clinic. The difference is who ends up in each box of the 2×2 table — and that depends entirely on which patients were enrolled in the validation study.

Study Population (Biased Spectrum): obvious disease vs. healthy controls
Severe Disease
Moderate
Mild
Borderline
Healthy Controls
Real Clinical Population (where the test must actually perform)
Severe
Moderate
Mild
Borderline / Ambiguous
True Negatives

In the biased study population, the test separates two easy groups — its performance looks excellent. In the real clinical population, most of the work is in the ambiguous middle, which the study never tested. Reported sensitivity and specificity collapse in practice.


Apply the Framework: Three Questions

Each scenario below describes a diagnostic test study. Apply the Part II framework to evaluate it. Track your progress below.

0 of 3 answered
Question 1 of 3 · Spectrum Bias
A new blood test for early-stage pancreatic cancer is evaluated. The study compares 80 patients with confirmed Stage III/IV pancreatic cancer to 80 healthy volunteers from a cancer screening program. Reported sensitivity: 94%. Which validity concern does this study most directly fail?
✓ Correct. This is a textbook spectrum bias scenario. The study pits the most obvious positive cases (advanced, confirmed cancer) against the most obvious negatives (healthy volunteers). The clinical challenge — distinguishing early-stage cancer from pancreatitis, cysts, or benign pancreatic changes — is entirely absent from the study population. The 94% sensitivity is likely to collapse when the test faces real clinical ambiguity.
Question 2 of 3 · Likelihood Ratios
A rapid antigen test for influenza has sensitivity 75% and specificity 95%. You're working in a walk-in clinic during peak flu season; your pre-test probability estimate for the patient in front of you is 40%. A test comes back positive. What is the approximate post-test probability of influenza?
LR+ = Sensitivity ÷ (1 − Specificity) = 0.75 ÷ 0.05 = 15. Pre-test odds = 0.40 ÷ 0.60 = 0.667. Post-test odds = 0.667 × 15 = 10.0. Post-test probability = 10 ÷ 11 ≈ ?
✓ Correct. Post-test probability = 10 ÷ (10 + 1) = 91%. A positive test with LR+ of 15 in a patient with moderate pre-test probability substantially rules in influenza. Note how different this is from the same test used in a low-prevalence setting: if pre-test probability were only 5%, post-test probability after a positive result would be only ~44% — essentially a coin flip. Same test, very different meaning. This is the Bayesian core of Part II.
Question 3 of 3 · Verification Bias
A study evaluates ultrasound for detecting appendicitis. Patients with a positive ultrasound are taken to surgery (and the appendix confirms or denies the diagnosis). Patients with a negative ultrasound are sent home and monitored — most are never re-evaluated by a gold standard. What bias does this introduce, and in which direction does it distort sensitivity?
✓ Correct. This is verification (work-up) bias. Because patients with negative ultrasounds are rarely operated on, the false negatives in cell c of the 2×2 table are systematically undercounted. Sensitivity = a ÷ (a + c), so when c is artificially small, sensitivity appears higher than it truly is. This is a pervasive problem in surgical diagnostic studies. Sackett's 1981 checklist made verification of all patients — regardless of index test result — a validity requirement for any diagnostic study claiming clinical relevance.
Final Score
0 / 3

The Diagnostic Test Framework Beyond Medicine

The bite-mark cases are the ones I return to most often when I think about Part II in a non-clinical context. The Innocence Project has documented approximately forty wrongful convictions based substantially on bite-mark analysis. Those forty people spent a collective 492 years in prison. The error rate for bite-mark identification has been reported as high as 91% in some studies — and as "low" as 11.9%. Either figure would be catastrophic for a diagnostic test. Neither figure was disclosed to most juries.

Apply the Part II checklist: What was the validation spectrum? Extreme cases. Was there an independent, blinded comparison to a gold standard? Often not — the forensic odontologist both collected and interpreted the evidence. Was verification bias present? Almost certainly — samples were only sent for analysis when investigators already had a suspect. Did the likelihood ratios account for the base rate? Never presented. The diagnostic framework isn't medicine. It's epistemology.

🦷 Bite-Mark Analysis
Sensitivity/Specificity: Error rates 11.9%–91% depending on study. No validated gold standard exists for human bite-mark uniqueness.

Spectrum bias: Studies used obvious match vs. obviously different samples. Real casework involves skin — elastic, post-mortem variable tissue.

Outcome: ~40 wrongful convictions, 492 collective years in prison. Nearly 25% of all exonerations since 1989 involved false forensic evidence.
🔬 Hair Microscopy
Sensitivity/Specificity: FBI review of 2,500 cases found microscopic hair analysis gave "erroneous testimony" in 90% of capital cases reviewed.

Verification bias: Hair was analyzed only after a suspect was identified — the analyst knew who they were "matching to."

Outcome: 143 cases reviewed in one NIJ study, with 59% involving errors contributing to wrongful conviction.
📰 Media Disease Scares
Classic failure: "Test that is 95% accurate" headlines ignore the base rate. A 95%-accurate test for a condition affecting 1 in 1,000 people produces 50 false positives for every true positive in a mass screening.

COVID example: Antigen tests with 98% specificity but 41% sensitivity in asymptomatic persons produced PPV of only 33% in low-prevalence settings — ignored by most media coverage in 2020.
🗳️ Political Polling
Pre-test probability matters: The "test" (poll) cannot be interpreted without knowing the prior probability of the event. A poll showing 52% support in a historically 70%–30% electorate is a very different signal than the same number in a genuinely competitive race.

Spectrum bias: Likely-voter screens validated on high-turnout elections fail in unusual-turnout election years — the validation spectrum doesn't match the deployment context.
Domain The "Diagnostic Test" Validity Question Most Often Failed Consequence
Criminal Justice Bite-mark, hair, serology analysis Spectrum bias + verification bias + blinding failure ~25% of wrongful convictions involve false forensic evidence
Public Health COVID rapid antigen screening Base rate / pre-test probability ignored in media reporting PPV 33% in asymptomatic low-prevalence setting
Cancer Screening Mammography (general population) Prevalence effect — 91% of positive mammograms in one model were false positives Unnecessary biopsies; 39%–44% of women who receive false positives do not return for future screening
Journalism Source credibility assessment Spectrum bias — journalists trained on obvious liars vs. obvious truth-tellers Difficulty detecting subtle, well-sourced propaganda
Intelligence Threat assessment indicators Verification bias — only confirmed threats are reviewed post-hoc Inflated apparent accuracy of profiling instruments

What Part III Adds

Part III — To Learn the Etiology or Causation of Disease — takes the analytical structure built in Parts I and II and applies it to a fundamentally harder problem: not "does this test work?" but "does this exposure cause this disease?" The leap from association to causation introduces a new set of concepts — relative risk, odds ratio, attributable risk, and the hierarchy of study designs for causal inference. The same discipline of structured critical appraisal, applied to a much trickier epistemological terrain.

← Part I: Why to Read Part III: Etiology → alexandrakitty.com