Clinical AI & Diagnostic Bias · Capstone brief
CLIN-07 · Imaging shortcut learning
An imaging model can post excellent numbers by reading something that is not the disease — a marker, a drain, a laterality token, a scanner signature, the fact that sicker patients get portable films. Take one documented shortcut, re-analyze or reproduce the evidence for it, and propose the check that would have caught it before anyone trusted the model.
The question
An imaging model can post excellent numbers by reading something that is not the disease — a marker, a drain, a laterality token, a scanner signature, the fact that sicker patients get portable films. Take one documented shortcut, re-analyze or reproduce the evidence for it, and propose the check that would have caught it before anyone trusted the model.
System / materials
Two viable routes; the teacher approves one in writing. Route A (no training required, always available): read a published shortcut finding and re-analyze its argument — what evidence would distinguish shortcut from signal, what the authors did, what a reviewer should have demanded. Route B: an open imaging dataset plus a small model, only where compute and access are genuinely available. Documented cases to start from: Zech et al., "Variable generalization performance of a deep learning model to detect pneumonia in chest radiographs," PLOS Medicine, 2018 — cross-site generalization collapse (https://journals.plos.org/plosmedicine/article?id=10.1371/journal.pmed.1002683); DeGrave et al., "AI for radiographic COVID-19 detection selects shortcuts over signal," Nature Machine Intelligence, 2021 (https://www.nature.com/articles/s42256-021-00338-7). No patient images from any placement, ever.
Expected failure modes
Producing a saliency map and calling it an explanation — saliency shows where, not why, and a plausible-looking heatmap has convinced many adults of nothing. Testing only on the source site, which is the condition under which every shortcut looks like skill. Treating a hypothesis about a shortcut as a demonstrated one. Route B attempted with neither compute nor access, ending in a half-trained model and no argument.
Done looks like
A shortcut analysis: the cue named concretely and the mechanism by which it correlates with the label; the evidence that separates shortcut from signal (external site, held-out scanner, cue removal or occlusion, or the published equivalent) with its limits stated; and a detection protocol — the specific check a team should run before deployment, written so another student could execute it.
Five C's
CT: distinguishing correlation with the label from the pathology itself. CR: designing a test that could actually falsify the shortcut hypothesis. CO: a peer proposes a rival explanation for the same evidence. CM: explaining the failure to a radiologic technologist without jargon. CZ: which patients are misread when the cue is absent or reversed.
Mentor role
A radiologist, radiologic technologist, or imaging researcher reviews the cue hypothesis and the proposed check. Standing instruction: reject saliency maps offered as proof. School-supervised.
Rubric calibration
R1: one cue, one modality, one claim. R2: the re-analysis or experiment is reproducible from the write-up. R3: comparator is same-site performance. R4: states what the evidence cannot rule out. R5: the detection protocol is executable by someone else. R6: refuses to present a hypothesis as a finding.
Two ways this goes wrong
(a) A heatmap gallery with no falsification test. (b) Route B chosen on optimism, abandoned at week five, with no analysis to submit.
Credit lane fit
Lane A immediately (health-science capstone; Route A also fits AP Research). No verified credit claim.