AI-Assisted Trades & Credentials · Capstone brief

TRD-01 · Diagnostic assistant limits

How often does it send a technician down a wrong path — a test, a teardown, or a part that costs hours and money before the error surfaces?

The question

Give an AI diagnostic assistant a set of scenarios whose real outcome is already known. How often does it send a technician down a wrong path — a test, a teardown, or a part that costs hours and money before the error surfaces? And for the answers it gets right, what did a technician still have to verify before touching anything?

Wrong-path rate is not the same as wrong-answer rate, and the difference is the brief. A model that reaches the correct fix after recommending an unnecessary component replacement has cost the shop a part, the labor to install it, and the customer's trust.

System / materials

Teacher names one trade and one symptom or task family in writing before Checkpoint 1 — the narrower the better. "Intermittent no-start on one engine family" or "no-cool call on one condenser type" produces a measurable result; "diagnostics" does not.

The student needs:

  • A scenario set with known ground truth — at least twelve, each with the actual root cause established. Sources in order of preference: completed repair orders from the school shop with customer, plate, and VIN details removed; instructor- or mentor-authored scenarios drawn from their own field experience; published training case studies or manufacturer diagnostic examples. Every scenario is written before any of it is shown to the model.
  • An AI tool the school permits, named with version in the log. Results do not transfer between tools or across versions, and the report must say so plainly.
  • An authority of record — the service manual, diagnostic tree, code book, or manufacturer procedure that establishes what the correct path was. The model is never the authority.
  • A cost basis — the shop's own labor guide or the instructor's flat-rate reference, plus parts pricing, so wrong paths convert into hours and dollars rather than adjectives.

Do not run a suggested diagnostic step on live customer property, a vehicle returning to service, or energized equipment. Scenarios are presented as text and, where the instructor approves, replicated on a trainer or scrap unit under the shop's normal sign-off.

Reference vocabulary (concepts — not a compliance claim):

  • The credential task list for the pathway, verified current per the credential rule above — for example the ASE Entry-Level task list for an automotive cohort (https://www.aseeducationfoundation.org/).
  • The shop's written diagnostic procedure, where one exists.

Expected failure modes

Scoring only the final answer and ignoring the route taken to it — the wrong-path cost disappears and the tool looks far better than it is. Feeding the model the symptom and hints from the known answer, which is the most common way this brief quietly invalidates itself. Judging the model's output as correct because it sounded like a technician. Counting a confident wrong answer the same as a hedged one, when confidence is exactly what determines whether a junior tech follows it. Scenarios written after seeing the model's responses. Twelve scenarios that are all the same fault wearing different clothes.

Done looks like

A tool characterization report with five pieces:

  1. Scenario set — twelve or more scenarios with established root causes, written and dated before any model contact, with identifying details removed and the source of each noted.
  2. Trial log — for each scenario: what the model was given (verbatim prompt), what it suggested as the first diagnostic step and full path, its stated or implied confidence, the actual root cause, and where the path diverged.
  3. Wrong-path accounting — the wrong-path rate, and each wrong path costed in labor hours and parts using the shop's own basis. Separate the categories that behave differently: unnecessary part replacement, unnecessary teardown, a missed test that would have caught it early, and a path that was merely inefficient rather than wrong.
  4. Required-verification list — for the suggestions the model got right, what a credentialed technician still had to confirm before acting. This list is the finding that TRD-04 turns into a checklist; state it as observations, not as a QC instrument.
  5. Refusal log — what the student concluded the tool must not be used for in this task family, with the evidence from the trial log that supports each item.

Customer, plate, VIN, and household details never appear in the scenario set or the report.

Five C's

CT: separating the answer from the route to it. CR: writing scenarios that could genuinely defeat the tool rather than flatter it. CO: a peer independently scores five trials for wrong-path versus inefficient-path; disagreements are reconciled and the scoring rule is rewritten until it holds. CM: a report a service manager would read before spending money on a subscription. CZ: who absorbs an unnecessary part — the customer, the warranty, or the shop.

Mentor role

A credentialed technician, shop foreman, or service manager reviews the scenario set at Checkpoint 1 — before any of it reaches the model — and the wrong-path costing at Checkpoint 2. Standing instruction: reject scenarios that leak the answer, and reject any plan touching live customer property. School-supervised, both times.

Rubric calibration

R1: one trade, one symptom family, twelve-plus scenarios, tool and version fixed. R2: verbatim prompts and responses logged; another student could replay the set. R3: comparator is the authority's diagnostic procedure — what the manual says the path should have been. R4: wrong paths categorized and costed in hours and dollars. R5: report is decision-ready for a service manager. R6: refusal log is specific and traceable to trials.

Two ways this goes wrong

(a) The student reports "the model was right 9 out of 12" with no path analysis, no cost, and no verification list — a percentage nobody can act on. (b) The scenarios are written after playing with the tool, so they unconsciously avoid the places it fails, and the report certifies a tool that has not been tested.

Checkpoint suggestions

  • Week 1–2: Trade and symptom family fixed; scenario set drafted, dated, and mentor-reviewed before any model contact; cost basis identified.
  • Week 4–5: Trials run and logged verbatim; peer scoring of five trials; scoring rule for wrong-path versus inefficient-path settled in writing.
  • Week 7–8: Wrong-path costing complete; required-verification list drafted; refusal log tied to specific trials.

Credit lane fit

Lane A immediately (center capstone, completer-year portfolio, or SkillsUSA project). Lane B with a division Internship wrapper and a real host shop. Documents practice toward credential competencies; it is not evidence of certification. No verified credit claim.