# PWR-193 — AI-assisted reasoning: learner worksheet

Release: Revision 7T · Full Tutorial Edition
Estimated time: Estimated 32 min reading and worksheet pass
Difficulty: Intermediate
Equipment: Common household or practice equipment — Declared AI-assisted reasoning fixture: A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations.; Setup aid for Solve before asking AI: Record the model version, settings and whether external tools or retrieval are enabled.; AI-assisted reasoning log: Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline; retain AI-assisted reasoning errors, assistance, stop and fallback.; AI-reasoning evidence card: use “Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task” to predeclare harmful acceptance, and “Humans inherit artificial intelligence biases” to track adopted error. Map the train-puzzle log to NIST AI RMF functions—measure, manage and document—without turning an explanation into verification.; Complete source sheet: R1 Red uses platform A on weekdays. R2 Red uses C on weekends. R3 Blue normally uses B. R4 during the listed Tuesday maintenance Blue moves to C. R5 Green uses B before 15:00 and A at or after 15:00. R6 Express uses C at odd-numbered hours and B at even-numbered hours. R7 Festival uses A on weekends and B on weekdays. R8 Shuttle uses C only when B is explicitly closed; otherwise it uses A. The rules are complete and the fictional week has Tuesday Blue maintenance only.; Supplied 12-item run, formatted ID|case|folded key|AI answer|AI premise note: 1|Red service on Monday|A|A|R1: weekday Red uses A. 2|Red service on Saturday|C|A|Incorrectly extends weekday R1. 3|Blue service on Wednesday with no maintenance|B|B|R3: normal Blue uses B. 4|Festival service on Tuesday|B|A|Incorrectly treats Tuesday as weekend under R7. 5|Green service at 14:00|B|B|R5: before 15:00 uses B. 6|Green service at 16:00|A|A|R5: at or after 15:00 uses A. 7|Express service at 11:00|C|C|R6: odd-hour Express uses C. 8|Express service at 12:00|B|C|Incorrectly applies the odd-hour clause. 9|Festival service on Sunday|A|B|Incorrectly applies the weekday clause. 10|Blue service on Tuesday during listed maintenance|C|C|R4 overrides normal Blue platform. 11|Shuttle on Friday while platform B is open|A|C|Invents a closure that is not on the sheet. 12|Red service on Sunday|C|B|Ignores weekend Red rule R2. Exactly six AI answers are correct and six are flawed. Keep the key column covered until human, AI and combined answers are locked.
Space: Desk / seated

Automated structural checklist: 10 of 10 structural checks present
Evidence quality/context: G1; Deep research depth
Automated method-quality band: Comprehensive
Method-rating basis: Structure coverage, instruction/check specificity, alternative distinctness, source troubleshooting, comparison depth and whether purpose/check fields are authored rather than derived.

Editorial review: Pending manual editorial sign off

## Before you begin
- [ ] I read the tutorial authority and stop conditions.
- [ ] I have the required equipment/space or a declared accessible alternative.

## Canonical source context

The train-scheduling pack supplies source rules and an AI that is right on six cases and wrong on six. Lock a human answer first, inspect each decisive premise, compare human-only, AI-only and combined scores, and retain harmful acceptance in the log. Improvement on these puzzles is evidence for a verification workflow, not superior reasoning everywhere.

Target outcome: On answer-key logic cases, the learner compares human-only, AI-only and combined answers, verifies each decisive premise and reports harmful acceptance.

### Authority and setup conditions

- Use a harmless answer-key task and exclude personal, proprietary or regulated data.
- Record the model version, settings and whether external tools or retrieval are enabled.
- Start check for AI-assisted reasoning: The run can be reproduced and contains no personal or consequential data.
- Top-of-sheet stop for AI-assisted reasoning: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.

### Required or supplied materials

- Declared AI-assisted reasoning fixture: A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations.
- Setup aid for Solve before asking AI: Record the model version, settings and whether external tools or retrieval are enabled.
- AI-assisted reasoning log: Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline; retain AI-assisted reasoning errors, assistance, stop and fallback.
- AI-reasoning evidence card: use “Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task” to predeclare harmful acceptance, and “Humans inherit artificial intelligence biases” to track adopted error. Map the train-puzzle log to NIST AI RMF functions—measure, manage and document—without turning an explanation into verification.
- Complete source sheet: R1 Red uses platform A on weekdays. R2 Red uses C on weekends. R3 Blue normally uses B. R4 during the listed Tuesday maintenance Blue moves to C. R5 Green uses B before 15:00 and A at or after 15:00. R6 Express uses C at odd-numbered hours and B at even-numbered hours. R7 Festival uses A on weekends and B on weekdays. R8 Shuttle uses C only when B is explicitly closed; otherwise it uses A. The rules are complete and the fictional week has Tuesday Blue maintenance only.
- Supplied 12-item run, formatted ID|case|folded key|AI answer|AI premise note: 1|Red service on Monday|A|A|R1: weekday Red uses A. 2|Red service on Saturday|C|A|Incorrectly extends weekday R1. 3|Blue service on Wednesday with no maintenance|B|B|R3: normal Blue uses B. 4|Festival service on Tuesday|B|A|Incorrectly treats Tuesday as weekend under R7. 5|Green service at 14:00|B|B|R5: before 15:00 uses B. 6|Green service at 16:00|A|A|R5: at or after 15:00 uses A. 7|Express service at 11:00|C|C|R6: odd-hour Express uses C. 8|Express service at 12:00|B|C|Incorrectly applies the odd-hour clause. 9|Festival service on Sunday|A|B|Incorrectly applies the weekday clause. 10|Blue service on Tuesday during listed maintenance|C|C|R4 overrides normal Blue platform. 11|Shuttle on Friday while platform B is open|A|C|Invents a closure that is not on the sheet. 12|Red service on Sunday|C|B|Ignores weekend Red rule R2. Exactly six AI answers are correct and six are flawed. Keep the key column covered until human, AI and combined answers are locked.
- Delayed no-AI strip: 13 Red on Thursday→A; 14 Express at 18:00→B; 15 Festival on Saturday→A; 16 Blue on Tuesday maintenance→C. Use the same source sheet, but no advice or prior answer may be visible.

## 1. Lock the task and version

Record the puzzle set, AI name/version, prompt, answer key, time limit and prohibited data before starting.

Why: The run can be reproduced and contains no personal or consequential data.

Success check: The run can be reproduced and contains no personal or consequential data.

Accessible alternative: Complete the same research action — “Record the puzzle set, AI name/version, prompt, answer key, time limit and prohibited data before starting.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “The run can be reproduced and contains no personal or consequential data.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 2. Solve before asking AI

Write an independent answer, premise chain and confidence for each item.

Why: Human-only performance is preserved before advice exposure.

Success check: Human-only performance is preserved before advice exposure.

Accessible alternative: Complete the same research action — “Write an independent answer, premise chain and confidence for each item.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “Human-only performance is preserved before advice exposure.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 3. Ask for a checkable decomposition

Request the AI to list premises, intermediate steps, conclusion and uncertainty rather than only an answer.

Why: Each conclusion can be traced to a stated rule.

Success check: Each conclusion can be traced to a stated rule.

Accessible alternative: A supported, seated, reduced-range or slower version of “Request the AI to list premises, intermediate steps, conclusion and uncertainty rather than only an answer.” is equivalent only when it keeps the same trained or measured target, declared configuration and success check: “Each conclusion can be traced to a stated rule.” If any of those change, record it as a related alternative rather than an equivalent repetition.

Alternative relationship: Conditional equivalence requires target and configuration check

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 4. Verify decisive premises

Open the source sheet and mark every cited departure rule true, false or absent.

Why: Unsupported and misquoted rules are visible before adoption.

Success check: Unsupported and misquoted rules are visible before adoption.

Accessible alternative: Complete “Open the source sheet and mark every cited departure rule true, false or absent.” in shorter passes, or use keyboard input, dictation or a support person, while preserving this success check: “Unsupported and misquoted rules are visible before adoption.” Use text-to-speech, structured premise tables or simplified displays without removing the independent-answer step.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 5. Search for a counterexample

Try one schedule that would make the AI conclusion fail; if found, reject or revise the reasoning.

Why: Accepted conclusions survive the declared counterexample test.

Success check: Accepted conclusions survive the declared counterexample test.

Accessible alternative: Complete “Try one schedule that would make the AI conclusion fail; if found, reject or revise the reasoning.” in shorter passes, or use keyboard input, dictation or a support person, while preserving this success check: “Accepted conclusions survive the declared counterexample test.” Allow calculators and extended time; compare each learner with their own matched conditions, not another person’s speed.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 6. Choose the final answer

Keep or change the human answer and write exactly which verified premise caused the decision.

Why: No change is justified by fluency or confidence alone.

Success check: No change is justified by fluency or confidence alone.

Accessible alternative: Complete the same research action — “Keep or change the human answer and write exactly which verified premise caused the decision.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “No change is justified by fluency or confidence alone.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 7. Score three conditions

Count accuracy, time and severity-weighted errors for human-only, AI-only and combined answers; also count harmful and helpful acceptance.

Why: Complementarity is claimed only if combined beats the better standalone baseline.

Success check: Complementarity is claimed only if combined beats the better standalone baseline.

Accessible alternative: Complete “Count accuracy, time and severity-weighted errors for human-only, AI-only and combined answers; also count harmful and helpful acceptance.” in shorter passes, or use keyboard input, dictation or a support person, while preserving this success check: “Complementarity is claimed only if combined beats the better standalone baseline.” Allow calculators and extended time; compare each learner with their own matched conditions, not another person’s speed.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 8. Retest without AI

After 48 hours solve four new items unaided, then compare with the original human baseline.

Why: Any device-on gain is separated from delayed unaided reasoning.

Success check: Any device-on gain is separated from delayed unaided reasoning.

Accessible alternative: Complete the same research action — “After 48 hours solve four new items unaided, then compare with the original human baseline.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “Any device-on gain is separated from delayed unaided reasoning.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## Reflection

What changed?

________________________________________________________________________________

What remains difficult?

________________________________________________________________________________

What will I repeat, adapt, ask for help with, or stop?

________________________________________________________________________________

---
Completion of this worksheet demonstrates tutorial participation only. It does not establish capability, qualification, safety clearance, diagnosis, treatment or independent validation.
