Episode II · CMAJ 1981 · Vol. 124 · pp. 703–710
To Learn the Features of
A Diagnostic Test
Sensitivity · Specificity · Likelihood Ratios · Spectrum Bias
An interactive exploration by Alexandra Kitty · alexandrakitty.com
Part II arrived about a week after Part I in our reading pile at McMaster, and I remember thinking the title was almost deceptively modest. "To learn the features of a diagnostic test." As if a test either works or it doesn't, and all we had to do was learn which was which.
What Sackett actually delivered was a controlled demolition of medical certainty. By the time I finished it, I understood that nearly every clinical test — including many that physicians ordered daily without a second thought — was operating in a probabilistic fog that most practitioners never acknowledged. The 2×2 table felt like a little bomb. How could something so simple expose so much?
The piece that stayed with me longest was the discussion of spectrum bias — the way a test validated on dramatic cases (obvious disease vs. obviously healthy) could completely collapse in the messy middle where it actually needed to work. I kept thinking about it years later when I started covering true crime cases. The "test" in those cases was often eyewitness testimony, or bite-mark analysis, or hair microscopy. And the "validation cohort" was almost never the population the test would face in a real investigation. They'd perfected the tool on perfect conditions, then deployed it in chaos.
The box in Hamilton is still packed. But I don't need to unpack it to reconstruct the argument — I've been applying it for forty-five years.
Part II of the How to Read Clinical Journals series appeared in the Canadian Medical Association Journal on March 15, 1981, in Volume 124, pages 703–710. Like Part I, it emerged from the McMaster University clinical epidemiology group that David Sackett was building into one of the most influential methodological research programs of the twentieth century. The article asked a deceptively simple question: when you read a paper about a diagnostic test, how do you decide whether the test is actually any good?
The answer, as with the rest of the series, was organized around a structured set of critical questions — questions that, once internalized, permanently change how you read any claim that a test "works."
1978 · The Paper That Lit the Fuse
Ransohoff & Feinstein: "Problems of Spectrum and Bias in Evaluating Diagnostic Tests"
Three years before Part II appeared, Alvan Feinstein and David Ransohoff published a landmark critique in the New England Journal of Medicine identifying two structural problems that undermined most published diagnostic test studies: spectrum bias and workup bias. Sackett built Part II directly on their framework, transforming it into a teachable set of appraisal criteria that a medical student — or a journalist, or a juror — could actually apply.
By 1981, diagnostic testing was proliferating. Biochemical analyzers were making it possible to run panels of 20 or 30 tests simultaneously on a single blood draw. Clinicians were ordering tests in industrial quantities, often without a clear sense of what a positive result actually meant in a patient with a moderate — not dramatic — probability of disease.
The central problem was Bayesian, though it was rarely framed that way. A test's ability to correctly identify disease (sensitivity) and correctly exclude disease (specificity) are properties of the test itself. But what a clinician actually needs to know — the probability that their specific patient has the disease, given a positive test result — depends critically on how common the disease is in that patient's context. The same test could be highly informative at a specialist referral centre and nearly useless in a primary care screening setting.
"A test with 95% sensitivity and 95% specificity will produce 50 false positives for every true positive when the disease prevalence is 1 in 1,000."
This counterintuitive arithmetic was not widely understood in 1981 — and, as COVID-19 rapid antigen testing controversies in 2020–2021 demonstrated, it remains poorly understood today. When COVID antigen tests showed 98–99% specificity but only 41% sensitivity in asymptomatic people, the positive predictive value in low-prevalence settings dropped to around 33%, meaning two-thirds of positive results in asymptomatic screening programs were false positives.
Sackett's Part II gave readers the tools to see this coming — four decades before the pandemic made it front-page news.
The Legacy Thread · 1981 → Today
From CMAJ to STARD, Cochrane, and the Fagan Nomogram
The structured appraisal framework in Part II directly seeded the QUADAS (Quality Assessment of Diagnostic Accuracy Studies) tool developed in 2003, and the STARD (Standards for Reporting of Diagnostic Accuracy) guidelines — now mandatory at leading medical journals. The "Users' Guide to the Medical Literature" chapter on diagnostic tests, which Sackett's McMaster group produced through JAMA in the 1990s, translated these principles to a new generation. Today, every systematic review of a diagnostic test conducted by the Cochrane Collaboration uses a checklist that traces its lineage back to the three questions in this 1981 paper.
The same test can appear near-perfect in a research study and near-useless in the clinic. The difference is who ends up in each box of the 2×2 table — and that depends entirely on which patients were enrolled in the validation study.
Study Population (Biased Spectrum): obvious disease vs. healthy controls
Severe Disease
Moderate
Mild
Borderline
Healthy Controls
Real Clinical Population (where the test must actually perform)
Severe
Moderate
Mild
Borderline / Ambiguous
True Negatives
In the biased study population, the test separates two easy groups — its performance looks excellent. In the real clinical population, most of the work is in the ambiguous middle, which the study never tested. Reported sensitivity and specificity collapse in practice.
Each scenario below describes a diagnostic test study. Apply the Part II framework to evaluate it. Track your progress below.
0 of 3 answered
Question 1 of 3 · Spectrum Bias
A new blood test for early-stage pancreatic cancer is evaluated. The study compares 80 patients with confirmed Stage III/IV pancreatic cancer to 80 healthy volunteers from a cancer screening program. Reported sensitivity: 94%. Which validity concern does this study most directly fail?
- A Verification bias — not all patients received the gold standard
- B Spectrum bias — the study used extreme cases instead of the clinically relevant ambiguous middle
- C Incorporation bias — the test result was used to define the gold standard
- D The study appears methodologically sound
✓ Correct. This is a textbook spectrum bias scenario. The study pits the most obvious positive cases (advanced, confirmed cancer) against the most obvious negatives (healthy volunteers). The clinical challenge — distinguishing early-stage cancer from pancreatitis, cysts, or benign pancreatic changes — is entirely absent from the study population. The 94% sensitivity is likely to collapse when the test faces real clinical ambiguity.
Question 2 of 3 · Likelihood Ratios
A rapid antigen test for influenza has sensitivity 75% and specificity 95%. You're working in a walk-in clinic during peak flu season; your pre-test probability estimate for the patient in front of you is 40%. A test comes back positive. What is the approximate post-test probability of influenza?
LR+ = Sensitivity ÷ (1 − Specificity) = 0.75 ÷ 0.05 = 15. Pre-test odds = 0.40 ÷ 0.60 = 0.667. Post-test odds = 0.667 × 15 = 10.0. Post-test probability = 10 ÷ 11 ≈ ?
- A About 75% — roughly equal to the test's sensitivity
- B About 60% — similar to the pre-test probability
- C About 91% — the positive test substantially revises probability upward
- D About 95% — equal to the test's specificity
✓ Correct. Post-test probability = 10 ÷ (10 + 1) = 91%. A positive test with LR+ of 15 in a patient with moderate pre-test probability substantially rules in influenza. Note how different this is from the same test used in a low-prevalence setting: if pre-test probability were only 5%, post-test probability after a positive result would be only ~44% — essentially a coin flip. Same test, very different meaning. This is the Bayesian core of Part II.
Question 3 of 3 · Verification Bias
A study evaluates ultrasound for detecting appendicitis. Patients with a positive ultrasound are taken to surgery (and the appendix confirms or denies the diagnosis). Patients with a negative ultrasound are sent home and monitored — most are never re-evaluated by a gold standard. What bias does this introduce, and in which direction does it distort sensitivity?
- A Spectrum bias — inflates specificity by including only severe cases
- B Verification bias — inflates sensitivity because false negatives are never confirmed
- C Incorporation bias — the ultrasound result was used to define the gold standard diagnosis
- D Review bias — surgeons who operated knew the ultrasound result
✓ Correct. This is verification (work-up) bias. Because patients with negative ultrasounds are rarely operated on, the false negatives in cell c of the 2×2 table are systematically undercounted. Sensitivity = a ÷ (a + c), so when c is artificially small, sensitivity appears higher than it truly is. This is a pervasive problem in surgical diagnostic studies. Sackett's 1981 checklist made verification of all patients — regardless of index test result — a validity requirement for any diagnostic study claiming clinical relevance.