Clinical AI & Diagnostic Bias · Capstone brief

CLIN-04 · Audit a published diagnostic model

Pick one published diagnostic or clinical risk model. Can you (1) reproduce or re-derive the headline performance the authors report, (2) convert it into positive and negative predictive values at a prevalence a clinician would actually see, (3) document where subgroup performance is unequal or simply absent, and (4) write the one-page caution sheet the paper did not — without claiming you validated anything?

The question

Pick one published diagnostic or clinical risk model. Can you (1) reproduce or re-derive the headline performance the authors report, (2) convert it into positive and negative predictive values at a prevalence a clinician would actually see, (3) document where subgroup performance is unequal or simply absent, and (4) write the one-page caution sheet the paper did not — without claiming you validated anything?

The hardest part of this brief is not the arithmetic. It is refusing to say more than the published evidence supports.

System / materials

Teacher approves one model and its public source in writing before Checkpoint 1. Three viable shapes, in ascending difficulty:

  • Metrics-only audit (safest, always available). A published paper, FDA summary, or model card that reports sensitivity, specificity, and cohort characteristics. The student re-derives predictive values, calibration implications, and subgroup gaps from the reported numbers alone. No dataset access required. This is a complete brief — do not treat it as the fallback.
  • Public-dataset re-analysis. The student re-runs a reported analysis on an openly downloadable dataset. Verify the download works from a school network before assigning.
  • Credentialed-dataset re-analysis. Only if the teacher holds credentialed access and supervises use under the data-use agreement. Students do not sign these agreements.

Candidate public sources to check availability on before Checkpoint 1 — all:

Reporting vocabulary (concepts — not a compliance claim):

Expected failure modes

Reporting AUC and stopping. Applying the paper's prevalence to a setting with a different one. Treating a missing subgroup breakdown as evidence of fairness — absence of a gap in the table is absence of a table. Quoting a number from an abstract without checking what population produced it. Writing "clinically validated" about a re-analysis. Slipping into clinical voice: recommending, diagnosing, or advising instead of auditing. Downloading a dataset whose access terms the student never read.

Done looks like

An audit memo packet with five pieces:

  1. Intended-use sentence — one sentence naming the model, the population, the clinical use (screening / triage / confirmation), and the decision it feeds. Copied from the source if the authors wrote one; written by the student and labeled as an inference if they did not.
  2. Metric reconstruction table — reported sensitivity and specificity, the prevalence the authors used, and PPV/NPV recomputed at two additional prevalences the student justifies (for example a high-prevalence specialty clinic and a general screening population). Show the arithmetic; a worked 2×2 beats a cited summary.
  3. Subgroup evidence map — for each subgroup the student can name (age, sex, race or ethnicity as reported, site, device, disease severity): reported / not reported / reported but underpowered. "Not reported" is the most common cell and must be visible in the table, not omitted.
  4. Caution sheet (one page) — plain language, preceptor-postable: what the model is for, what it must not be used for, the two populations where performance is least supported, and the failure that would hurt someone. No hedging adverbs, no hype.
  5. Refusal log — what the student declined to claim: no validation claim, no deployment recommendation, no PHI, no inference about a subgroup the source never measured.

Raw downloads and any credentialed data stay on school-managed storage. The showcase version carries the memo, the tables, and the caution sheet — nothing else.

Five C's

CT: converting a reported metric into a predictive value at a named prevalence. CR: choosing the two comparison prevalences and defending them. CO: a peer independently recomputes one PPV cell and the two disagree until they find the error. CM: a caution sheet a preceptor reads once and posts. CZ: naming who absorbs a false negative and who absorbs a false positive — separately, by name of role or population.

Mentor role

A clinician, clinical informaticist, laboratory scientist, or public-health analyst reviews the intended-use sentence and prevalence choices at Checkpoint 1, and the caution sheet at Checkpoint 2. Standing instruction: reject any draft that reads as clinical advice, and reject any plan that touches PHI. School-supervised, both times.

Rubric calibration

R1: one model, one population, intended-use sentence written. R2: every number traceable to a cited public source with version and access route. R3: prevalence-only or existing-clinical-rule comparator present. R4: subgroup map includes "not reported" rows; base-rate effect shown numerically. R5: caution sheet fits one page and a preceptor could post it. R6: refusal log names the specific claims declined, not "I was ethical."

Two ways this goes wrong

(a) The student reproduces the abstract, adds a chart, and calls it an audit — no prevalence conversion, no missing-evidence table, nothing a clinician could act on. (b) The student finds a subgroup gap and overclaims it as proof of discrimination, when the sample in that cell was forty patients and the confidence interval spans the whole range.

Checkpoint suggestions

  • Week 1–2: Model and public source approved in writing; access route confirmed to actually work; intended-use sentence drafted; PHI rule signed.
  • Week 4–5: Metric reconstruction table with the two justified prevalences; peer recomputation of one cell; subgroup evidence map first pass, with "not reported" rows already filled in.
  • Week 7–8: Caution sheet reviewed by the mentor; refusal log complete; showcase version stripped of any raw data.

Credit lane fit

Lane A immediately (health-science senior capstone, HOSA project, or AP Research). Lane B only with a division Internship wrapper and a real host — and the host placement still does not permit PHI in the capstone artifact. No verified credit claim.