Clinical AI & Diagnostic Bias · Capstone brief
CLIN-02 · Sensitivity is a choice
Take one model with public scores or a published ROC, rebuild the confusion matrix at two operating points, and defend which point belongs in a screening use and which belongs in a confirmatory use. Whose error did you choose to accept?
The question
Sensitivity and specificity are not properties of a model — they are a pair you selected when you picked a threshold. Take one model with public scores or a published ROC, rebuild the confusion matrix at two operating points, and defend which point belongs in a screening use and which belongs in a confirmatory use. Whose error did you choose to accept?
System / materials
One of: a published ROC curve or multi-threshold table the student can read points off (approve the source's resolution before assigning — a low-resolution figure will not support two clean points); or a public dataset where the student can compute continuous scores from a simple model and choose thresholds directly. Teacher fixes the prevalence in writing so it does not float between the two comparisons. Reporting shape: TRIPOD+AI's discrimination and calibration headings (https://www.equator-network.org/reporting-guidelines/tripod-statement/).
Expected failure modes
Reporting AUC as if it settled the threshold question — AUC is the whole curve, and no patient is ever tested at "the whole curve." Picking Youden's J because a tutorial said so, with no clinical reason. Letting prevalence change between the two points. Calling the high-sensitivity point "safer" without counting the follow-up burden it creates.
Done looks like
A threshold memo: both confusion matrices side by side at fixed prevalence; sensitivity, specificity, PPV, NPV, and raw false-positive and false-negative counts per point; a paragraph naming the downstream consequence of each error type in this specific clinical path (what test, what wait, what cost, what harm); and a recommendation of which point fits screening and which fits confirmation, with the reasoning stated in harms rather than in metrics.
Five C's
CT: reading a threshold as a value judgment. CR: choosing the two points with a clinical rationale. CO: a peer argues the opposite recommendation from the same tables. CM: the harms paragraph, in plain language. CZ: naming who bears the missed cases and who bears the false alarms.
Mentor role
A clinician or laboratory scientist reviews the two chosen operating points and the harms paragraph. Standing instruction: reject any memo that recommends a threshold without naming the follow-up pathway it triggers. School-supervised.
Rubric calibration
R1: one model, one prevalence, two named points. R2: both matrices reproducible from cited numbers. R3: comparator is the other operating point, argued honestly. R4: both error types counted as people. R5: harms paragraph readable by a non-specialist. R6: refuses to declare one threshold universally correct.
Two ways this goes wrong
(a) Two thresholds, four metrics, no mention of what happens to the person who tests positive. (b) Prevalence quietly changes between the two tables, so the PPV comparison is meaningless.
Credit lane fit
Lane A immediately (health-science seminar or a statistics cross-listing). No verified credit claim.