Clinical AI & Diagnostic Bias · Capstone brief

CLIN-01 · 95% accurate and still wrong

Holding the operating point completely fixed, can you show what happens to its positive predictive value as prevalence moves across the settings that test would actually be used in — and then write the explanation a patient or a first-year nursing student would understand on the first read?

The question

A test or model is announced at "95% accurate." Holding the operating point completely fixed, can you show what happens to its positive predictive value as prevalence moves across the settings that test would actually be used in — and then write the explanation a patient or a first-year nursing student would understand on the first read?

Nothing here requires a computer beyond a calculator. That is the point: the failure this brief targets is not a modeling failure, it is a reading failure, and clinicians and journalists make it constantly.

System / materials

Teacher approves one test or model with publicly reported sensitivity and specificity before Checkpoint 1. The student never chooses the numbers; the source does.

Acceptable sources:

  • A screening or diagnostic test with characteristics published by a public health body or professional guideline.
  • A published clinical prediction model that reports sensitivity and specificity at a stated threshold.
  • A regulatory summary or model card reporting the same two numbers.

If the source reports only "accuracy," the brief cannot start. Accuracy alone does not determine a 2×2 without prevalence. Sending the student back to find sensitivity and specificity — or to a different source — is the first real lesson, not a delay.

The teacher also fixes the three prevalences in writing: one from the source's own cohort, and two the student justifies from public reporting (for example a symptomatic specialty clinic and a general asymptomatic screening population). Prevalences are cited, never invented.

Reference vocabulary (concepts — not a clinical claim):

Expected failure modes

Reading PPV off the sensitivity number — the single most common error, and the one the brief exists to break. Changing the threshold partway through, which turns this into CLIN-02 and destroys the comparison. Inventing a prevalence because the real one was hard to find. Reporting PPV to four decimal places from a cohort of two hundred. Writing the explainer in textbook voice, so it explains nothing to the person who needs it. Concluding "the test is bad" — a low PPV in a low-prevalence setting is usually the test working exactly as designed in the wrong place.

Done looks like

A numeracy packet with five pieces:

  1. Source card — the test or model, its cited sensitivity and specificity, the threshold they were measured at, the cohort that produced them, and the link. One paragraph.
  2. Three worked 2×2 tables — same sensitivity, same specificity, same threshold; three cited prevalences. Built by hand from a stated population of 10,000 so every cell is a whole number of people. PPV and NPV computed and shown, not asserted.
  3. Natural-frequency paragraph — the headline result restated as "out of 10,000 people like this, X test positive, of whom Y actually have the condition." No percentages in this paragraph at all.
  4. The "what accuracy hides" table — one row per prevalence: overall accuracy, PPV, NPV, and the count of false positives. The row where accuracy stays high while PPV collapses is the deliverable; mark it.
  5. Refusal log — no threshold changes, no invented prevalence, no claim about whether the test should be used, no extrapolation to a population the source never studied.

Hand-built tables are acceptable and often better. A spreadsheet is fine. A model is not required and earns no extra credit here.

Five C's

CT: separating a conditional probability from its reverse. CR: choosing and defending the two comparison prevalences. CO: a peer recomputes one full 2×2 independently; the pair reconciles any disagreement before either table is final. CM: the natural-frequency paragraph, tested on one reader outside the class who reports back what they understood. CZ: who receives the false positives — the follow-up biopsies, the anxiety, the cost — and who receives the false negatives.

Mentor role

A clinician, nurse educator, laboratory scientist, or public-health analyst reviews the prevalence choices at Checkpoint 1 and the natural-frequency paragraph at Checkpoint 2 — reading it as the intended audience would, not as an expert. Standing instruction: reject any draft that recommends for or against using the test. School-supervised, both times.

Rubric calibration

R1: one test, one fixed threshold, three cited prevalences. R2: every cell reproducible by hand from the stated numbers. R3: comparator is the prevalence-only guess — what you would conclude with no test at all. R4: false-positive counts shown as people, and the small-cohort limits of the source stated. R5: the natural-frequency paragraph survives contact with a non-expert reader. R6: refusal log names the specific claims declined.

Two ways this goes wrong

(a) The student computes three PPVs correctly, writes them as a percentage table, and never converts to people — so the finding lands on nobody. (b) The student concludes the test is worthless, when the actual finding is that it is being read wrong in a population it was never validated for.

Checkpoint suggestions

  • Week 1–2: Source card approved; sensitivity, specificity, and threshold frozen in writing; the three prevalences cited.
  • Week 3–4: All three 2×2 tables built by hand; peer recomputation of one table; disagreements resolved on paper.
  • Week 5–6: Natural-frequency paragraph tested on one outside reader and revised; "what accuracy hides" table finalized; refusal log complete.

Credit lane fit

Lane A immediately (health-science seminar, HOSA project, or a statistics or AP Research cross-listing). Runs in a shorter window than the rest of the bank — six weeks is realistic. No verified credit claim.