Adversarial ML & AI Security · Capstone brief

SEC-12 · Defend your own prior work

Take a model or assistant you already built (Applied ML brief, a class project, or a public baseline the teacher names) and apply one mitigation from SEC-01 or SEC-04’s defense menu. Does the original success metric survive, and does the attack-class success rate drop in the same sandbox rules as this track?

← All briefsIn the bank PDF · use Print

School-approved sandbox only. No production targets. Showcase materials paraphrase class behavior — do not publish payload strings.

The question

Take a model or assistant you already built (Applied ML brief, a class project, or a public baseline the teacher names) and apply one mitigation from SEC-01 or SEC-04’s defense menu. Does the original success metric survive, and does the attack-class success rate drop in the same sandbox rules as this track?

Lab / materials

Student’s prior artifact plus the same approved sandbox constraints as SEC-01/04. If the student has no prior model, teacher assigns a public baseline (tiny classifier or school sandbox assistant). Do not “defend” a production district system.

Expected failure modes

Adding a disclaimer instead of a control. Breaking the original task and calling it security. Re-running a new attack class without a baseline from earlier work.

Done looks like

Before/after on (a) original task metric and (b) the chosen class’s lab success rate, plus a short model-card addendum: new limits, remaining residual risk, refusal log.

Five C's

CT: security that destroys the product is not done. CR: one mitigation, measured. CO: a peer tries the same class against your defended artifact under lab rules. CM: addendum a later student could inherit. CZ: shipping undefended classwork into a showcase without limits.

Mentor role

Same as SEC-04 — reviews the mitigation choice at Checkpoint 1 and the before/after at Checkpoint 2. School-supervised.

Rubric calibration

R1: one prior artifact, one class, one mitigation. R2: both metrics regenerable. R3: undefended prior work is the baseline. R4: residual risk named. R5: addendum is short. R6: still sandbox-only.

Two ways this goes wrong

(a) “We added an ethics paragraph.” (b) The model now refuses every useful question and the attack rate is “zero.”

Source moved or something unclear? Send feedback on this brief. Mentors are advisory; the school supervises. These briefs do not produce verified computer-science credit.