Assistive & Accessible AI · Capstone brief

ACC-11 · Eval with real tasks

Replace demo scripts with tasks users actually need; report completion, errors, and abandonment.

The question

Replace demo scripts with tasks users actually need; report completion, errors, and abandonment.

Partners / materials

Consented partners or published task analyses — no invented users. Task set from real needs (communicate X, navigate Y, complete Z). Metrics: completion, errors, time, abandonment — declared before sessions. Demo script explicitly banned as the graded path.

Expected failure modes

Happy-path demo only. Ignoring abandonment. Averaging away the partner who could not finish. Claiming statistical power from three people.

Done looks like

Task-based eval report with counts; abandonment narratives (consented detail level); comparison to the old demo script; recommendation to ship, revise, or kill.

Five C's

CT: task success vs demo theater. CR: real tasks. CO: partner or peer confirms tasks are real. CM: eval report. CZ: users who abandon.

Mentor role

Accessibility mentor reviews task authenticity. School-supervised.

Rubric calibration

R1: one product, real tasks. R2: metrics pre-declared. R3: demo script as negative control. R4: abandonment recorded. R5: report decisive. R6: sample limits stated.

Two ways this goes wrong

(a) Scripted wow-demo as “eval.” (b) Percentages from three partners sold as proof.

Credit lane fit

Lane A immediately. No verified credit claim.