Assistive & Accessible AI · Capstone brief

ACC-06 · Speech tech accents and dialects

Measure ASR error across dialects/accents using public data or consented samples; propose mitigation.

The question

Measure ASR error across dialects/accents using public data or consented samples; propose mitigation.

Partners / materials

Public speech datasets or consented recordings under school ethics rules — never scraped social audio, never secret recordings. Error metric declared before listening (e.g. WER on a fixed transcript set). Mitigation proposals that do not require the speaker to “sound more standard.”

Expected failure modes

One accent as “the problem.” Tiny samples treated as population truth. Recommending that users change their speech. Inventing demographic labels for speakers.

Done looks like

Methods note; error table with counts; sample-limit sentence; mitigation options with owners; non-claim list about populations not measured.

Five C's

CT: error as system failure, not speaker failure. CR: fixed transcript set. CO: peer re-scores two clips. CM: mitigation memo. CZ: speakers mis-recognized by default models.

Mentor role

Linguistics, ESL, or special-ed teacher reviews framing. Standing instruction: reject “speak clearer” as the primary fix. School-supervised.

Rubric calibration

R1: one ASR system, defined speaker set. R2: metric and clips documented. R3: baseline condition named. R4: sample limits stated. R5: memo actionable. R6: non-claim list.

Two ways this goes wrong

(a) “Accents don’t work” with n=2. (b) Mitigation that blames the speaker.

Credit lane fit

Lane A immediately. No verified credit claim.