Assistive & Accessible AI · Capstone brief
ACC-06 · Speech tech accents and dialects
Measure ASR error across dialects/accents using public data or consented samples; propose mitigation.
The question
Measure ASR error across dialects/accents using public data or consented samples; propose mitigation.
Partners / materials
Public speech datasets or consented recordings under school ethics rules — never scraped social audio, never secret recordings. Error metric declared before listening (e.g. WER on a fixed transcript set). Mitigation proposals that do not require the speaker to “sound more standard.”
Expected failure modes
One accent as “the problem.” Tiny samples treated as population truth. Recommending that users change their speech. Inventing demographic labels for speakers.
Done looks like
Methods note; error table with counts; sample-limit sentence; mitigation options with owners; non-claim list about populations not measured.
Five C's
CT: error as system failure, not speaker failure. CR: fixed transcript set. CO: peer re-scores two clips. CM: mitigation memo. CZ: speakers mis-recognized by default models.
Mentor role
Linguistics, ESL, or special-ed teacher reviews framing. Standing instruction: reject “speak clearer” as the primary fix. School-supervised.
Rubric calibration
R1: one ASR system, defined speaker set. R2: metric and clips documented. R3: baseline condition named. R4: sample limits stated. R5: memo actionable. R6: non-claim list.
Two ways this goes wrong
(a) “Accents don’t work” with n=2. (b) Mitigation that blames the speaker.
Credit lane fit
Lane A immediately. No verified credit claim.