Machines & Meaning · Capstone brief

HUM-03 · What tokenization erases

Show a meaning-bearing distinction that breaks under BPE/WordPiece — and why it matters for a text you care about.

The question

Show a meaning-bearing distinction that breaks under BPE/WordPiece — and why it matters for a text you care about.

Texts / materials

One literary, historical, or multilingual passage the student cares about. Tokenizer inspection (school-approved library or demo). Before/after token table. Short argument tying the break to interpretation — without claiming the model “misunderstands” as a mind.

Expected failure modes

Random weird tokens with no interpretive stake. Declaring tokenization evil in general. Inventing linguistics papers. Skipping the actual text.

Done looks like

Passage citation; token table; interpretive memo; non-claim list (this does not prove machines lack meaning; it shows a representation limit for this distinction).

Five C's

CT: representation vs meaning-talk. CR: text-tied example. CO: peer finds another break in the same passage. CM: memo. CZ: readers of the text whose distinction vanishes.

Mentor role

Language/literature teacher reviews interpretive stake. School-supervised.

Rubric calibration

R1: one passage, one distinction. R2: tokens shown. R3: human reading baseline. R4: scope limited to this case. R5: memo clear. R6: non-claim list on mind/meaning.

Two ways this goes wrong

(a) Token dump, no stakes. (b) “Therefore LLMs don’t understand language.”

Credit lane fit

Lane A immediately. No verified credit claim.