Section 379 of 440
Complete canonical tutorial. This reader section contains the same teaching body as PWR-197 · AI-assisted decision calibration. Open the Power dossier.
PWR-197 · SELF GUIDED full tutorial
Accept and reject AI advice in proportion to its demonstrated reliability
The travel-budget deck hides individual answer keys while exposing HIGH or LOW AI confidence. Recalculate every total independently, then record helpful acceptance, harmful acceptance and missed help by band before checking the key. The observed 8/10 versus 3/10 reliability belongs to this supplied deck and cannot set policy for another model or decision.
1 · Permission and limits
Know exactly what you may do
2 · Get ready
Gather what you need and check the starting conditions
What you need
- Declared AI-assisted decision calibration fixture: A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong.
- Setup aid for Make and lock a first choice: Do not use real travel bookings, money movement or personal data.
- AI-assisted decision calibration log: Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata; retain AI-assisted decision calibration errors, assistance, stop and fallback.
- Calibration source sheet: use “Humans inherit artificial intelligence biases” and “Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task” to list when people adopt wrong output and whether explanations prevent it. Predeclare harmful acceptance and missed help for HIGH and LOW budget advice. Use NIST AI RMF to record model version, tested deck and limits instead of treating confidence wording as reliability.
- Budget source rule: every fictional case total is nights×£20 + the stated rail fare + meal units×£5. No tax, discount or unstated fee applies. A final answer within £0 is correct; accepting wrong advice is harmful and rejecting correct advice is missed help.
- Complete 20-case deck, formatted ID|case|folded key|AI advice|confidence band: 1|1 night, rail £10, 2 meal units|40|40|HIGH. 2|2 nights, rail £15, 2 meal units|65|65|HIGH. 3|1 night, rail £5, 3 meal units|40|40|HIGH. 4|3 nights, rail £12, 1 meal unit|77|77|HIGH. 5|2 nights, rail £8, 4 meal units|68|68|HIGH. 6|1 night, rail £12, 2 meal units|42|57|HIGH. 7|3 nights, rail £20, 2 meal units|90|90|HIGH. 8|1 night, rail £9, 1 meal unit|34|34|HIGH. 9|1 night, rail £6, 1 meal unit|31|31|LOW. 10|2 nights, rail £14, 3 meal units|69|69|HIGH. 11|4 nights, rail £10, 2 meal units|100|105|LOW. 12|1 night, rail £17, 2 meal units|47|52|HIGH. 13|2 nights, rail £11, 1 meal unit|56|56|LOW. 14|3 nights, rail £7, 3 meal units|82|87|LOW. 15|1 night, rail £13, 4 meal units|53|58|LOW. 16|2 nights, rail £19, 2 meal units|69|74|LOW. 17|3 nights, rail £9, 1 meal unit|74|74|LOW. 18|1 night, rail £21, 2 meal units|51|56|LOW. 19|2 nights, rail £5, 3 meal units|60|65|LOW. 20|4 nights, rail £12, 1 meal unit|97|102|LOW. High-band advice is correct on 8/10 (all except 6 and 12). low-band advice is correct on 3/10 (9, 13 and 17). Keep the key folded until every first answer, AI exposure, independent check and final choice is recorded.
- Five delayed no-AI cases, formatted ID|case|folded key: 21|2 nights, rail £7, 2 meal units|57; 22|1 night, rail £16, 3 meal units|51; 23|3 nights, rail £11, 2 meal units|81; 24|2 nights, rail £18, 1 meal unit|63; 25|1 night, rail £4, 4 meal units|44. Do not show advice or earlier error notes during this check.
Before you start
- Use an answer-key task with deliberately mixed advice quality.
- Do not use real travel bookings, money movement or personal data.
- Start check for AI-assisted decision calibration: The calibration rule is written before advice appears.
- Top-of-sheet stop for AI-assisted decision calibration: Stop if advice affects real money, health, employment, law, benefits or safety.
3 · The method
Follow these steps in order
- Define the decision and loss
State the harmless budget target, acceptable error and what counts as helpful or harmful acceptance.
Why: The calibration rule is written before advice appears.
Check: The calibration rule is written before advice appears.
- Make and lock a first choice
Solve each case and record confidence 0–100 before revealing AI advice.
Why: Human baseline and uncertainty are preserved.
Check: Human baseline and uncertainty are preserved.
- Reveal advice with metadata
Log AI answer, confidence band, version and any explanation without yet opening the key.
Why: Advice quality can be stratified.
Check: Advice quality can be stratified; verify it in the travel-budget advice deck.
- Apply a forcing check
Before accepting, cite one independent arithmetic or source check and write one reason the advice could be wrong.
Why: Acceptance is delayed until an external check exists.
Check: Acceptance is delayed until an external check exists.
- Choose accept, reject or defer
Make the final choice and record which check controlled it.
Why: The choice is not explained by confidence alone.
Check: The choice is not explained by confidence alone.
- Score reliance outcomes
With the key, count correct advice accepted, wrong advice rejected, wrong advice accepted and correct advice rejected.
Why: Appropriate reliance, over-reliance and under-reliance are separate.
Check: Appropriate reliance, over-reliance and under-reliance are separate.
- Plot calibration bands
For each human and AI confidence band, compare stated confidence with actual accuracy.
Why: A 90-confidence band is not called calibrated if it is right 60% of the time.
Check: A 90-confidence band is not called calibrated if it is right 60% of the time.
- Test persistence of bias
After 48 hours, solve five similar cases without AI and check whether previously wrong advice appears in answers.
Why: Acquired AI bias is measured rather than assumed away.
Check: Acquired AI bias is measured rather than assumed away.
4 · Worked example
See the whole method used once
Scenario
A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong.
Walkthrough
- Solve case 6 as £42 at 65 confidence before seeing AI advice.
- The AI says £57 at high confidence; independently recalculate the supplied line items and find an invented £15 surcharge, so reject the advice.
- For case 9, a source-table check confirms the AI’s low-confidence £31 answer, so accept despite the low band.
- After 20 cases, count 7 helpful acceptances, 5 justified rejections, 2 harmful acceptances and 1 missed helpful answer among disagreements.
- Compare the high-confidence band’s 8/10 accuracy with the low band’s 3/10 instead of using a single trust score.
- Two days later, solve five new cases unaided and flag any repeated double-tax error as possible acquired bias.
Result
The learner calibrates by tested band and independent checks, while two harmful acceptances remain visible. The deck does not set a safe reliance rule for another task or model.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
| Moment | Right / safer | Wrong / riskier | Why it matters |
|---|---|---|---|
| Forcing check | Do independent arithmetic before acceptance when completing the travel-budget advice deck. | Read the explanation and treat it as verification. | Explanations may not reduce harm from wrong advice. |
| High confidence | Use demonstrated band accuracy plus case evidence. | Accept every 90-confidence answer during the travel-budget advice deck. | Confidence can be miscalibrated; this distorts the travel-budget advice deck. |
| Reliance metric | Count harmful and missed-helpful advice separately. | Report only final task accuracy during the travel-budget advice deck. | The same accuracy can hide very different reliance errors. |
| Delayed bias | Test new unaided cases after advice exposure. | Assume rejected wrong advice leaves no trace. | AI bias can persist into later human decisions. |
6 · Common mistakes
Spot the error and apply the correction
| Mistake | Fix |
|---|---|
| Read the explanation and treat it as verification. | Require a source or calculation outside the model response. |
| Accept every 90-confidence answer during the travel-budget advice deck. | Score each confidence band against the key. |
| Report only final task accuracy. | Fill all four reliance cells within the travel-budget advice deck. |
| Assume rejected wrong advice leaves no trace. | Include a delayed no-AI check within the travel-budget advice deck. |
7 · Practice
Turn the steps into a usable skill
First session
- Travel-budget calibration run: A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong.
- Predeclared loss and first total: Define the decision and loss: State the harmless budget target, acceptable error and what counts as helpful or harmful acceptance.
- Independent arithmetic acceptance check: Apply a forcing check, then Choose accept, reject or defer.
- Explanation-as-proof correction: if “Read the explanation and treat it as verification.” appears, apply “Require a source or calculation outside the model response.”
- Five-case bias check: Test persistence of bias: After 48 hours, solve five similar cases without AI and check whether previously wrong advice appears in answers.
Repeat plan
Complete one 20-case deck, then five no-AI items after 48 hours. Repeat monthly with changed reliability bands and version labels; discard the old rule after any task, model or distribution change.
Progress when
- The calibration rule is written before advice appears.
- Human baseline and uncertainty are preserved.
- Advice quality can be stratified.
- Harmful acceptance decreases across matched decks, helpful acceptance is retained, confidence bands match observed accuracy within the predeclared tolerance, and delayed bias does not increase.
Do not progress when
- Do not continue while this error remains: Read the explanation and treat it as verification.
- Pause until this correction works: Score each confidence band against the key.
- This AI-assisted decision calibration stop ends the block: Stop if advice affects real money, health, employment, law, benefits or safety.
8 · Check the result
Measure what changed
Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata
How: Baseline fixture: A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong. Enter success checks from “Define the decision and loss” and “Choose accept, reject or defer”. If “Read the explanation and treat it as verification.” occurs, apply its named fix; then score Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata on the unused “Test persistence of bias” item.
Good result: Harmful acceptance decreases across matched decks, helpful acceptance is retained, confidence bands match observed accuracy within the predeclared tolerance, and delayed bias does not increase.
This does not prove: Boundary for AI-assisted decision calibration: “Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata” describes only A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong. It cannot establish “A confidence score makes AI safe to trust”.
Self-check
- Without the example, demonstrate: The calibration rule is written before advice appears.
- Find the fault in this attempt: “Read the explanation and treat it as verification.” Apply “Require a source or calculation outside the model response.”; what changes?
- What evidence in the completed record shows that this is wrong: “Accept every 90-confidence answer.”?
- AI-assisted decision calibration stop decision: Stop if advice affects real money, health, employment, law, benefits or safety.
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
- Stop if advice affects real money, health, employment, law, benefits or safety.
- Do not reuse a calibration rate after a model, prompt, population or task change.
- Pause if independent checking is unavailable; defer rather than accept on confidence.
Accessibility and adaptations
- Use a three-level confidence scale with observed proportions written beside each level.
- Provide calculator, enlarged tables or extra time; speed is reported as workload, not calibration quality.
10 · Evidence and limits
Why these instructions are here
- primary research
Registered support for AI-assisted decision calibration: “Humans inherit artificial intelligence biases”. It bears on Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata inside the AI-assisted decision calibration fixture. It does not validate “A confidence score makes AI safe to trust”.
Humans inherit artificial intelligence biases - primary research
Constraint for AI-assisted decision calibration, drawn from “Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task”: Incorrect AI reduced accuracy, explanations often failed to help, and acquired AI bias persisted into later unaided decisions.
Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task - official guidance
NIST AI-risk application to AI-assisted decision calibration: declare the system, preserve rights, verify outputs and log failure. The local test is “Choose accept, reject or defer”; its registered observation is Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata.
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Limits
- AI-assisted decision calibration boundary: interpret “Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata” only for A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong.
- A successful result does not establish “A confidence score makes AI safe to trust”.
- AI-assisted decision calibration limiting finding: Incorrect AI reduced accuracy, explanations often failed to help, and acquired AI bias persisted into later unaided decisions.
- No perfect-performance claim for AI-assisted decision calibration: the evidence register does not make “Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata” universal, consequence-free or flawless in A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong.
- Scope remains AI-assisted decision calibration: A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong. Recheck the comparator, support and “Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata” after any configuration change.
Open the complete canonical research register
- Primary empirical supportLimiting / contraryHumans inherit artificial intelligence biases
Lucía Vicente; Helena Matute · 2023 · Primary research
- Primary empirical supportLimiting / contraryExplainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task
Julia Cecil; Eva Lermer; Matthias F. C. Hudecek; Jan Sauer; Susanne Gaube · 2024 · Primary research
- Primary empirical supportLimiting / contraryDoes the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance
Gagan Bansal; Tongshuang Wu; Joyce Zhou; Raymond Fok; Besmira Nushi; Ece Kamar; Marco Tulio Ribeiro; Daniel S. Weld · 2021 · Primary research
- Primary empirical supportLimiting / contraryTo Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
Zana Buçinca; Maja Barbara Malaya; Krzysztof Z. Gajos · 2021 · Primary research
- Primary empirical supportLimiting / contraryThe impact of AI errors in a human-in-the-loop process
Ujué Agudo; Karlos G. Liberal; Miren Arrese; Helena Matute · 2024 · Primary research
- Primary empirical supportLimiting / contraryEffect of Uncertainty-Aware AI Models on Pharmacists' Reaction Time and Decision-Making in a Web-Based Mock Medication Verification Task: Randomized Controlled Trial
Corey Lester; Brigid Rowell; Yifan Zheng; Zoe Co; Vincent Marshall; Jin Yong Kim; Qiyuan Chen; Raed Kontar; X. Jessie Yang · 2025 · Primary research
- Limiting / contraryOfficial boundary contextArtificial Intelligence Risk Management Framework (AI RMF 1.0)
Elham Tabassi; National Institute of Standards and Technology · 2023 · Official standard
- Limiting / contraryOfficial boundary contextNIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0
National Institute of Standards and Technology · 2020 · Official standard
Read the complete evidence interpretation on the Power dossier.
Tutorial delivery controls
Learn, adapt, troubleshoot and resume
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
Define the decision and loss
State the harmless budget target, acceptable error and what counts as helpful or harmful acceptance.
The calibration rule is written before advice appears.
The calibration rule is written before advice appears.
I’m stuck on this step
Reset: Re-read this authored instruction — “State the harmless budget target, acceptable error and what counts as helpful or harmful acceptance.” — and its success check, then attempt only this step.
Possible snag: The result from “State the harmless budget target, acceptable error and what counts as helpful or harmful acceptance.” does not yet meet this declared check: The calibration rule is written before advice appears.
Correction: Return to the start of “Define the decision and loss”, reduce complexity or pace, and repeat only the part needed to satisfy: “The calibration rule is written before advice appears.”
Stop / get help: Stop if advice affects real money, health, employment, law, benefits or safety.
Make and lock a first choice
Solve each case and record confidence 0–100 before revealing AI advice.
Human baseline and uncertainty are preserved.
Human baseline and uncertainty are preserved.
I’m stuck on this step
Reset: Re-read this authored instruction — “Solve each case and record confidence 0–100 before revealing AI advice.” — and its success check, then attempt only this step.
Possible snag: The result from “Solve each case and record confidence 0–100 before revealing AI advice.” does not yet meet this declared check: Human baseline and uncertainty are preserved.
Correction: Return to the start of “Make and lock a first choice”, reduce complexity or pace, and repeat only the part needed to satisfy: “Human baseline and uncertainty are preserved.”
Stop / get help: Stop if advice affects real money, health, employment, law, benefits or safety.
Reveal advice with metadata
Log AI answer, confidence band, version and any explanation without yet opening the key.
Advice quality can be stratified.
Advice quality can be stratified; verify it in the travel-budget advice deck.
I’m stuck on this step
Reset: Re-read this authored instruction — “Log AI answer, confidence band, version and any explanation without yet opening the key.” — and its success check, then attempt only this step.
Possible snag: Read the explanation and treat it as verification.
Correction: Require a source or calculation outside the model response.
Possible snag: Accept every 90-confidence answer during the travel-budget advice deck.
Correction: Score each confidence band against the key.
Stop / get help: Stop if advice affects real money, health, employment, law, benefits or safety.
Apply a forcing check
Before accepting, cite one independent arithmetic or source check and write one reason the advice could be wrong.
Acceptance is delayed until an external check exists.
Acceptance is delayed until an external check exists.
I’m stuck on this step
Reset: Re-read this authored instruction — “Before accepting, cite one independent arithmetic or source check and write one reason the advice could be wrong.” — and its success check, then attempt only this step.
Possible snag: Assume rejected wrong advice leaves no trace.
Correction: Include a delayed no-AI check within the travel-budget advice deck.
Stop / get help: Stop if advice affects real money, health, employment, law, benefits or safety.
Choose accept, reject or defer
Make the final choice and record which check controlled it.
The choice is not explained by confidence alone.
The choice is not explained by confidence alone.
I’m stuck on this step
Reset: Re-read this authored instruction — “Make the final choice and record which check controlled it.” — and its success check, then attempt only this step.
Possible snag: The result from “Make the final choice and record which check controlled it.” does not yet meet this declared check: The choice is not explained by confidence alone.
Correction: Return to the start of “Choose accept, reject or defer”, reduce complexity or pace, and repeat only the part needed to satisfy: “The choice is not explained by confidence alone.”
Stop / get help: Stop if advice affects real money, health, employment, law, benefits or safety.
Score reliance outcomes
With the key, count correct advice accepted, wrong advice rejected, wrong advice accepted and correct advice rejected.
Appropriate reliance, over-reliance and under-reliance are separate.
Appropriate reliance, over-reliance and under-reliance are separate.
I’m stuck on this step
Reset: Re-read this authored instruction — “With the key, count correct advice accepted, wrong advice rejected, wrong advice accepted and correct advice rejected.” — and its success check, then attempt only this step.
Possible snag: Report only final task accuracy.
Correction: Fill all four reliance cells within the travel-budget advice deck.
Stop / get help: Stop if advice affects real money, health, employment, law, benefits or safety.
Plot calibration bands
For each human and AI confidence band, compare stated confidence with actual accuracy.
A 90-confidence band is not called calibrated if it is right 60% of the time.
A 90-confidence band is not called calibrated if it is right 60% of the time.
I’m stuck on this step
Reset: Re-read this authored instruction — “For each human and AI confidence band, compare stated confidence with actual accuracy.” — and its success check, then attempt only this step.
Possible snag: The result from “For each human and AI confidence band, compare stated confidence with actual accuracy.” does not yet meet this declared check: A 90-confidence band is not called calibrated if it is right 60% of the time.
Correction: Return to the start of “Plot calibration bands”, reduce complexity or pace, and repeat only the part needed to satisfy: “A 90-confidence band is not called calibrated if it is right 60% of the time.”
Stop / get help: Stop if advice affects real money, health, employment, law, benefits or safety.
Test persistence of bias
After 48 hours, solve five similar cases without AI and check whether previously wrong advice appears in answers.
Acquired AI bias is measured rather than assumed away.
Acquired AI bias is measured rather than assumed away.
I’m stuck on this step
Reset: Re-read this authored instruction — “After 48 hours, solve five similar cases without AI and check whether previously wrong advice appears in answers.” — and its success check, then attempt only this step.
Possible snag: The result from “After 48 hours, solve five similar cases without AI and check whether previously wrong advice appears in answers.” does not yet meet this declared check: Acquired AI bias is measured rather than assumed away.
Correction: Return to the start of “Test persistence of bias”, reduce complexity or pace, and repeat only the part needed to satisfy: “Acquired AI bias is measured rather than assumed away.”
Stop / get help: Stop if advice affects real money, health, employment, law, benefits or safety.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Do independent arithmetic before acceptance when completing the travel-budget advice deck.
Read the explanation and treat it as verification.
Use demonstrated band accuracy plus case evidence.
Accept every 90-confidence answer during the travel-budget advice deck.
Count harmful and missed-helpful advice separately.
Report only final task accuracy during the travel-budget advice deck.
Test new unaided cases after advice exposure.
Assume rejected wrong advice leaves no trace.
Method-structure checklist
10 of 10 structural checks present
- Ordered, Power-specific instructions — present
- Every activity has a success check — present
- Materials or supplied records are declared — present
- Measurement or assessment rule is present — present
- Tutorial-specific troubleshooting is present — present
- Stopping or escalation boundary is present — present
- Every activity has an adjacent alternative — present
- Correct-versus-incorrect comparison is present — present
- Evidence context is bound to the Power record — present
- Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.