Clinical AI & Diagnostic Bias · Capstone brief
CLIN-12 · Decide not to deploy
Then evaluate a model that meets its headline target and fails your bar, and write the recommendation against deployment. Can you hold a line you drew before you knew what it would cost you?
The question
Set your bar in week one — fairness, calibration, subgroup coverage, whatever you can defend — in writing, before you see results. Then evaluate a model that meets its headline target and fails your bar, and write the recommendation against deployment. Can you hold a line you drew before you knew what it would cost you?
System / materials
Any model from earlier in this bank, a classmate's, or a published model with sufficient reported detail. The distinguishing requirement is the pre-registration: the bar is written, dated, and countersigned by the teacher in week one, and it does not change afterward. If the model unexpectedly clears the bar, that is a legitimate outcome — the deliverable becomes the deployment recommendation with its conditions and monitoring plan, and the reasoning is graded identically.
Expected failure modes
Setting the bar after seeing results, which is the whole failure this brief targets. Setting a bar so lax nothing could fail it, or so strict nothing could ever pass — both are ways of avoiding the decision. Recommending against deployment with no account of what patients lose by not deploying. Confusing "I do not like this model" with an argued case.
Done looks like
A decision memo: the dated pre-registered bar with the reasoning for each threshold; the evaluation against it, pass and fail marked per criterion; the cost of the no — what the status quo continues to cost, named honestly; the recommendation with its conditions for revisiting; and a paragraph on what would have to be true for the answer to change.
Five C's
CT: holding a pre-specified criterion against a result you wanted. CR: setting a bar that could plausibly go either way. CO: a peer or mentor argues the deployment case and is answered on the merits. CM: a memo a committee could act on. CZ: harmed by deploying, and harmed by waiting — both, named.
Mentor role
A clinician or informaticist countersigns the bar at Checkpoint 1 — before any evaluation — and reviews the memo at Checkpoint 2. Standing instruction: reject any bar submitted after results exist. School-supervised.
Rubric calibration
R1: bar written and dated in week one. R2: evaluation reproducible against each stated criterion. R3: comparator is the status quo, costed. R4: both directions of harm quantified as far as evidence allows. R5: memo is committee-ready. R6: this brief is R6 — the refusal is the artifact.
Two ways this goes wrong
(a) The bar is quietly rewritten in week six to match what the model turned out to do. (b) A confident no with no acknowledgment of what the no costs the patients who would have been helped.
Credit lane fit
Lane A immediately, and the natural capstone finale for this track. Pairs with CLIN-04 or CLIN-06 as the evaluation input. No verified credit claim.