Section 375 of 440
Complete canonical tutorial. This reader section contains the same teaching body as PWR-193 · AI-assisted reasoning. Open the Power dossier.
PWR-193 · SELF GUIDED full tutorial
Use AI as a fallible second reasoner while preserving an independent answer and source check
The train-scheduling pack supplies source rules and an AI that is right on six cases and wrong on six. Lock a human answer first, inspect each decisive premise, compare human-only, AI-only and combined scores, and retain harmful acceptance in the log. Improvement on these puzzles is evidence for a verification workflow, not superior reasoning everywhere.
1 · Permission and limits
Know exactly what you may do
2 · Get ready
Gather what you need and check the starting conditions
What you need
- Declared AI-assisted reasoning fixture: A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations.
- Setup aid for Solve before asking AI: Record the model version, settings and whether external tools or retrieval are enabled.
- AI-assisted reasoning log: Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline; retain AI-assisted reasoning errors, assistance, stop and fallback.
- AI-reasoning evidence card: use “Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task” to predeclare harmful acceptance, and “Humans inherit artificial intelligence biases” to track adopted error. Map the train-puzzle log to NIST AI RMF functions—measure, manage and document—without turning an explanation into verification.
- Complete source sheet: R1 Red uses platform A on weekdays. R2 Red uses C on weekends. R3 Blue normally uses B. R4 during the listed Tuesday maintenance Blue moves to C. R5 Green uses B before 15:00 and A at or after 15:00. R6 Express uses C at odd-numbered hours and B at even-numbered hours. R7 Festival uses A on weekends and B on weekdays. R8 Shuttle uses C only when B is explicitly closed; otherwise it uses A. The rules are complete and the fictional week has Tuesday Blue maintenance only.
- Supplied 12-item run, formatted ID|case|folded key|AI answer|AI premise note: 1|Red service on Monday|A|A|R1: weekday Red uses A. 2|Red service on Saturday|C|A|Incorrectly extends weekday R1. 3|Blue service on Wednesday with no maintenance|B|B|R3: normal Blue uses B. 4|Festival service on Tuesday|B|A|Incorrectly treats Tuesday as weekend under R7. 5|Green service at 14:00|B|B|R5: before 15:00 uses B. 6|Green service at 16:00|A|A|R5: at or after 15:00 uses A. 7|Express service at 11:00|C|C|R6: odd-hour Express uses C. 8|Express service at 12:00|B|C|Incorrectly applies the odd-hour clause. 9|Festival service on Sunday|A|B|Incorrectly applies the weekday clause. 10|Blue service on Tuesday during listed maintenance|C|C|R4 overrides normal Blue platform. 11|Shuttle on Friday while platform B is open|A|C|Invents a closure that is not on the sheet. 12|Red service on Sunday|C|B|Ignores weekend Red rule R2. Exactly six AI answers are correct and six are flawed. Keep the key column covered until human, AI and combined answers are locked.
- Delayed no-AI strip: 13 Red on Thursday→A; 14 Express at 18:00→B; 15 Festival on Saturday→A; 16 Blue on Tuesday maintenance→C. Use the same source sheet, but no advice or prior answer may be visible.
Before you start
- Use a harmless answer-key task and exclude personal, proprietary or regulated data.
- Record the model version, settings and whether external tools or retrieval are enabled.
- Start check for AI-assisted reasoning: The run can be reproduced and contains no personal or consequential data.
- Top-of-sheet stop for AI-assisted reasoning: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
3 · The method
Follow these steps in order
- Lock the task and version
Record the puzzle set, AI name/version, prompt, answer key, time limit and prohibited data before starting.
Why: The run can be reproduced and contains no personal or consequential data.
Check: The run can be reproduced and contains no personal or consequential data.
- Solve before asking AI
Write an independent answer, premise chain and confidence for each item.
Why: Human-only performance is preserved before advice exposure.
Check: Human-only performance is preserved before advice exposure.
- Ask for a checkable decomposition
Request the AI to list premises, intermediate steps, conclusion and uncertainty rather than only an answer.
Why: Each conclusion can be traced to a stated rule.
Check: Each conclusion can be traced to a stated rule.
- Verify decisive premises
Open the source sheet and mark every cited departure rule true, false or absent.
Why: Unsupported and misquoted rules are visible before adoption.
Check: Unsupported and misquoted rules are visible before adoption.
- Search for a counterexample
Try one schedule that would make the AI conclusion fail; if found, reject or revise the reasoning.
Why: Accepted conclusions survive the declared counterexample test.
Check: Accepted conclusions survive the declared counterexample test.
- Choose the final answer
Keep or change the human answer and write exactly which verified premise caused the decision.
Why: No change is justified by fluency or confidence alone.
Check: No change is justified by fluency or confidence alone.
- Score three conditions
Count accuracy, time and severity-weighted errors for human-only, AI-only and combined answers; also count harmful and helpful acceptance.
Why: Complementarity is claimed only if combined beats the better standalone baseline.
Check: Complementarity is claimed only if combined beats the better standalone baseline.
- Retest without AI
After 48 hours solve four new items unaided, then compare with the original human baseline.
Why: Any device-on gain is separated from delayed unaided reasoning.
Check: Any device-on gain is separated from delayed unaided reasoning.
4 · Worked example
See the whole method used once
Scenario
A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations.
Walkthrough
- Record the puzzle set, model version and exact prompt; solve item 4 as “Platform B” with 70 confidence.
- The AI says “Platform A” and cites rule 7, so list its premise chain before changing anything.
- Check the source sheet and find that rule 7 applies only on weekends; this case is Tuesday.
- Build a Tuesday schedule that satisfies every rule and reaches Platform B, rejecting the AI answer.
- Across 12 items, score human-only 8, AI-only 6 and combined 10, with two harmful suggestions rejected and one harmful suggestion accepted.
- Two days later, solve four unused items unaided and report 3/4 without claiming transfer beyond this puzzle family.
Result
The combined process beats both standalone scores in this fixed puzzle set while retaining one harmful acceptance. It demonstrates a verification workflow, not general superior reasoning.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
| Moment | Right / safer | Wrong / riskier | Why it matters |
|---|---|---|---|
| Independent baseline | Solve and lock an answer before AI exposure. | Ask AI first and reconstruct a “human” answer later. | The comparison becomes contaminated and early-answer influence is hidden. |
| Premise verification | Check each decisive premise against the source sheet. | Accept an explanation because it is detailed. | Fluent explanations can contain false rules. |
| Counterexample in train-platform reasoning deck | Try to falsify the conclusion with one valid schedule. | Ask the same model whether it agrees with itself. | Self-confirmation is not independent verification; this distorts the train-platform reasoning deck. |
| Complementarity in train-platform reasoning deck | Compare combined performance with the better of human-only and AI-only. | Call any combined improvement over human alone a success. | The AI alone may already be better, or bad advice may lower performance. |
6 · Common mistakes
Spot the error and apply the correction
| Mistake | Fix |
|---|---|
| Ask AI first and reconstruct a “human” answer later. | Timestamp the first answer and confidence. |
| Accept an explanation because it is detailed. | Mark each premise true, false or absent. |
| Ask the same model whether it agrees with itself. | Use the answer key or source constraints. |
| Call any combined improvement over human alone a success. | Report all three baselines and harm-weighted errors. |
7 · Practice
Turn the steps into a usable skill
First session
- Train-platform reasoning run: A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations.
- Locked human premise chain: Lock the task and version: Record the puzzle set, AI name/version, prompt, answer key, time limit and prohibited data before starting.
- Source and counterexample audit: Verify decisive premises, then Search for a counterexample.
- AI-first contamination correction: if “Ask AI first and reconstruct a “human” answer later.” appears, apply “Timestamp the first answer and confidence.”
- Four-item no-AI delay: Retest without AI: After 48 hours solve four new items unaided, then compare with the original human baseline.
Repeat plan
Use 8–12 answer-key items once per week. Retest four unseen items without AI after 48 hours and change the AI error mix monthly; progress only when combined beats the better standalone baseline without more severe errors.
Progress when
- The run can be reproduced and contains no personal or consequential data.
- Human-only performance is preserved before advice exposure.
- Each conclusion can be traced to a stated rule.
- On two new sets, every adopted AI premise is source-verified, harmful acceptance is no higher than one item, and combined accuracy exceeds both declared standalone conditions.
Do not progress when
- Do not continue while this error remains: Ask AI first and reconstruct a “human” answer later.
- Pause until this correction works: Mark each premise true, false or absent.
- This AI-assisted reasoning stop ends the block: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
8 · Check the result
Measure what changed
Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline
How: Baseline fixture: A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations. Enter success checks from “Lock the task and version” and “Search for a counterexample”. If “Ask AI first and reconstruct a “human” answer later.” occurs, apply its named fix; then score Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline on the unused “Retest without AI” item.
Good result: On two new sets, every adopted AI premise is source-verified, harmful acceptance is no higher than one item, and combined accuracy exceeds both declared standalone conditions.
This does not prove: Boundary for AI-assisted reasoning: “Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline” describes only A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations. It cannot establish “AI makes a person universally smarter”.
Self-check
- Without the example, demonstrate: The run can be reproduced and contains no personal or consequential data.
- Find the fault in this attempt: “Ask AI first and reconstruct a “human” answer later.” Apply “Timestamp the first answer and confidence.”; what changes?
- What evidence in the completed record shows that this is wrong: “Accept an explanation because it is detailed.”?
- AI-assisted reasoning stop decision: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
- Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
- Do not paste confidential or personal data into an unapproved system.
- Suspend the run after unexplained model, source or tool changes until a new baseline is made.
Accessibility and adaptations
- Use text-to-speech, structured premise tables or simplified displays without removing the independent-answer step.
- Allow calculators and extended time; compare each learner with their own matched conditions, not another person’s speed.
10 · Evidence and limits
Why these instructions are here
- primary research
Registered support for AI-assisted reasoning: “Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task”. It bears on Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline inside the AI-assisted reasoning fixture. It does not validate “AI makes a person universally smarter”.
Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task - primary research
Constraint for AI-assisted reasoning, drawn from “Humans inherit artificial intelligence biases”: Explanations did not reliably create complementarity and incorrect advice sometimes reduced performance below the human-only baseline.
Humans inherit artificial intelligence biases - official guidance
NIST AI-risk application to AI-assisted reasoning: declare the system, preserve rights, verify outputs and log failure. The local test is “Search for a counterexample”; its registered observation is Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline.
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Limits
- AI-assisted reasoning boundary: interpret “Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline” only for A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations.
- A successful result does not establish “AI makes a person universally smarter”.
- AI-assisted reasoning limiting finding: Explanations did not reliably create complementarity and incorrect advice sometimes reduced performance below the human-only baseline.
- No perfect-performance claim for AI-assisted reasoning: the evidence register does not make “Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline” universal, consequence-free or flawless in A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations.
- Scope remains AI-assisted reasoning: A 12-item fictional train-scheduling puzzle has a source sheet of departure rules. A declared AI system gives six correct and six deliberately flawed explanations. Recheck the comparator, support and “Accuracy, calibration, time, verification effort and harm-weighted error versus human-only, AI-only and the better standalone baseline” after any configuration change.
Open the complete canonical research register
- Limiting / contraryHumans inherit artificial intelligence biases
Lucía Vicente; Helena Matute · 2023 · Primary research
- Primary empirical supportLimiting / contraryExplainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task
Julia Cecil; Eva Lermer; Matthias F. C. Hudecek; Jan Sauer; Susanne Gaube · 2024 · Primary research
- Primary empirical supportLimiting / contraryExperimental evidence on the productivity effects of generative artificial intelligence
Shakked Noy; Whitney Zhang · 2023 · Primary research
- Primary empirical supportLimiting / contraryDoes the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
Gagan Bansal; Tongshuang Wu; Joyce Zhou; Raymond Fok; Besmira Nushi; Ece Kamar; Marco Tulio Ribeiro; Daniel S. Weld · 2021 · Primary research
- Primary empirical supportLimiting / contraryTo Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
Zana Buçinca; Maja Barbara Malaya; Krzysztof Z. Gajos · 2021 · Primary research
- Limiting / contraryThe impact of AI errors in a human-in-the-loop process
Ujué Agudo; Karlos G. Liberal; Miren Arrese; Helena Matute · 2024 · Primary research
- Limiting / contraryEffect of Uncertainty-Aware AI Models on Pharmacists' Reaction Time and Decision-Making in a Web-Based Mock Medication Verification Task: Randomized Controlled Trial
Corey Lester; Brigid Rowell; Yifan Zheng; Zoe Co; Vincent Marshall; Jin Yong Kim; Qiyuan Chen; Raed Kontar; X. Jessie Yang · 2025 · Primary research
- Limiting / contraryOfficial boundary contextArtificial Intelligence Risk Management Framework (AI RMF 1.0)
Elham Tabassi; National Institute of Standards and Technology · 2023 · Official standard
- Limiting / contraryOfficial boundary contextArtificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
C. Autio; R. Schwartz; J. Dunietz; S. Jain; M. Stanley; E. Tabassi; P. Hall; National Institute of Standards and Technology · 2024 · Official standard
Read the complete evidence interpretation on the Power dossier.
Tutorial delivery controls
Learn, adapt, troubleshoot and resume
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
Lock the task and version
Record the puzzle set, AI name/version, prompt, answer key, time limit and prohibited data before starting.
The run can be reproduced and contains no personal or consequential data.
The run can be reproduced and contains no personal or consequential data.
I’m stuck on this step
Reset: Re-read this authored instruction — “Record the puzzle set, AI name/version, prompt, answer key, time limit and prohibited data before starting.” — and its success check, then attempt only this step.
Possible snag: Ask the same model whether it agrees with itself.
Correction: Use the answer key or source constraints.
Stop / get help: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
Solve before asking AI
Write an independent answer, premise chain and confidence for each item.
Human-only performance is preserved before advice exposure.
Human-only performance is preserved before advice exposure.
I’m stuck on this step
Reset: Re-read this authored instruction — “Write an independent answer, premise chain and confidence for each item.” — and its success check, then attempt only this step.
Possible snag: Ask AI first and reconstruct a “human” answer later.
Correction: Timestamp the first answer and confidence.
Stop / get help: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
Ask for a checkable decomposition
Request the AI to list premises, intermediate steps, conclusion and uncertainty rather than only an answer.
Each conclusion can be traced to a stated rule.
Each conclusion can be traced to a stated rule.
I’m stuck on this step
Reset: Re-read this authored instruction — “Request the AI to list premises, intermediate steps, conclusion and uncertainty rather than only an answer.” — and its success check, then attempt only this step.
Possible snag: The result from “Request the AI to list premises, intermediate steps, conclusion and uncertainty rather than only an answer.” does not yet meet this declared check: Each conclusion can be traced to a stated rule.
Correction: Return to the start of “Ask for a checkable decomposition”, reduce complexity or pace, and repeat only the part needed to satisfy: “Each conclusion can be traced to a stated rule.”
Stop / get help: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
Verify decisive premises
Open the source sheet and mark every cited departure rule true, false or absent.
Unsupported and misquoted rules are visible before adoption.
Unsupported and misquoted rules are visible before adoption.
I’m stuck on this step
Reset: Re-read this authored instruction — “Open the source sheet and mark every cited departure rule true, false or absent.” — and its success check, then attempt only this step.
Possible snag: Accept an explanation because it is detailed.
Correction: Mark each premise true, false or absent.
Stop / get help: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
Search for a counterexample
Try one schedule that would make the AI conclusion fail; if found, reject or revise the reasoning.
Accepted conclusions survive the declared counterexample test.
Accepted conclusions survive the declared counterexample test.
I’m stuck on this step
Reset: Re-read this authored instruction — “Try one schedule that would make the AI conclusion fail; if found, reject or revise the reasoning.” — and its success check, then attempt only this step.
Possible snag: The result from “Try one schedule that would make the AI conclusion fail; if found, reject or revise the reasoning.” does not yet meet this declared check: Accepted conclusions survive the declared counterexample test.
Correction: Return to the start of “Search for a counterexample”, reduce complexity or pace, and repeat only the part needed to satisfy: “Accepted conclusions survive the declared counterexample test.”
Stop / get help: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
Choose the final answer
Keep or change the human answer and write exactly which verified premise caused the decision.
No change is justified by fluency or confidence alone.
No change is justified by fluency or confidence alone.
I’m stuck on this step
Reset: Re-read this authored instruction — “Keep or change the human answer and write exactly which verified premise caused the decision.” — and its success check, then attempt only this step.
Possible snag: The result from “Keep or change the human answer and write exactly which verified premise caused the decision.” does not yet meet this declared check: No change is justified by fluency or confidence alone.
Correction: Return to the start of “Choose the final answer”, reduce complexity or pace, and repeat only the part needed to satisfy: “No change is justified by fluency or confidence alone.”
Stop / get help: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
Score three conditions
Count accuracy, time and severity-weighted errors for human-only, AI-only and combined answers; also count harmful and helpful acceptance.
Complementarity is claimed only if combined beats the better standalone baseline.
Complementarity is claimed only if combined beats the better standalone baseline.
I’m stuck on this step
Reset: Re-read this authored instruction — “Count accuracy, time and severity-weighted errors for human-only, AI-only and combined answers; also count harmful and helpful acceptance.” — and its success check, then attempt only this step.
Possible snag: Call any combined improvement over human alone a success.
Correction: Report all three baselines and harm-weighted errors.
Stop / get help: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
Retest without AI
After 48 hours solve four new items unaided, then compare with the original human baseline.
Any device-on gain is separated from delayed unaided reasoning.
Any device-on gain is separated from delayed unaided reasoning.
I’m stuck on this step
Reset: Re-read this authored instruction — “After 48 hours solve four new items unaided, then compare with the original human baseline.” — and its success check, then attempt only this step.
Possible snag: The result from “After 48 hours solve four new items unaided, then compare with the original human baseline.” does not yet meet this declared check: Any device-on gain is separated from delayed unaided reasoning.
Correction: Return to the start of “Retest without AI”, reduce complexity or pace, and repeat only the part needed to satisfy: “Any device-on gain is separated from delayed unaided reasoning.”
Stop / get help: Stop if the task becomes medical, legal, financial, hiring, security or safety critical.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Solve and lock an answer before AI exposure.
Ask AI first and reconstruct a “human” answer later.
Check each decisive premise against the source sheet.
Accept an explanation because it is detailed.
Try to falsify the conclusion with one valid schedule.
Ask the same model whether it agrees with itself.
Compare combined performance with the better of human-only and AI-only.
Call any combined improvement over human alone a success.
Method-structure checklist
10 of 10 structural checks present
- Ordered, Power-specific instructions — present
- Every activity has a success check — present
- Materials or supplied records are declared — present
- Measurement or assessment rule is present — present
- Tutorial-specific troubleshooting is present — present
- Stopping or escalation boundary is present — present
- Every activity has an adjacent alternative — present
- Correct-versus-incorrect comparison is present — present
- Evidence context is bound to the Power record — present
- Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.