# PWR-197 — AI-assisted decision calibration: learner worksheet

Release: Revision 7T · Full Tutorial Edition
Estimated time: Estimated 32 min reading and worksheet pass
Difficulty: Intermediate
Equipment: Common household or practice equipment — Declared AI-assisted decision calibration fixture: A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong.; Setup aid for Make and lock a first choice: Do not use real travel bookings, money movement or personal data.; AI-assisted decision calibration log: Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata; retain AI-assisted decision calibration errors, assistance, stop and fallback.; Calibration source sheet: use “Humans inherit artificial intelligence biases” and “Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task” to list when people adopt wrong output and whether explanations prevent it. Predeclare harmful acceptance and missed help for HIGH and LOW budget advice. Use NIST AI RMF to record model version, tested deck and limits instead of treating confidence wording as reliability.; Budget source rule: every fictional case total is nights×£20 + the stated rail fare + meal units×£5. No tax, discount or unstated fee applies. A final answer within £0 is correct; accepting wrong advice is harmful and rejecting correct advice is missed help.; Complete 20-case deck, formatted ID|case|folded key|AI advice|confidence band: 1|1 night, rail £10, 2 meal units|40|40|HIGH. 2|2 nights, rail £15, 2 meal units|65|65|HIGH. 3|1 night, rail £5, 3 meal units|40|40|HIGH. 4|3 nights, rail £12, 1 meal unit|77|77|HIGH. 5|2 nights, rail £8, 4 meal units|68|68|HIGH. 6|1 night, rail £12, 2 meal units|42|57|HIGH. 7|3 nights, rail £20, 2 meal units|90|90|HIGH. 8|1 night, rail £9, 1 meal unit|34|34|HIGH. 9|1 night, rail £6, 1 meal unit|31|31|LOW. 10|2 nights, rail £14, 3 meal units|69|69|HIGH. 11|4 nights, rail £10, 2 meal units|100|105|LOW. 12|1 night, rail £17, 2 meal units|47|52|HIGH. 13|2 nights, rail £11, 1 meal unit|56|56|LOW. 14|3 nights, rail £7, 3 meal units|82|87|LOW. 15|1 night, rail £13, 4 meal units|53|58|LOW. 16|2 nights, rail £19, 2 meal units|69|74|LOW. 17|3 nights, rail £9, 1 meal unit|74|74|LOW. 18|1 night, rail £21, 2 meal units|51|56|LOW. 19|2 nights, rail £5, 3 meal units|60|65|LOW. 20|4 nights, rail £12, 1 meal unit|97|102|LOW. High-band advice is correct on 8/10 (all except 6 and 12). low-band advice is correct on 3/10 (9, 13 and 17). Keep the key folded until every first answer, AI exposure, independent check and final choice is recorded.
Space: Desk / seated

Automated structural checklist: 10 of 10 structural checks present
Evidence quality/context: G1; Deep research depth
Automated method-quality band: Comprehensive
Method-rating basis: Structure coverage, instruction/check specificity, alternative distinctness, source troubleshooting, comparison depth and whether purpose/check fields are authored rather than derived.

Editorial review: Pending manual editorial sign off

## Before you begin
- [ ] I read the tutorial authority and stop conditions.
- [ ] I have the required equipment/space or a declared accessible alternative.

## Canonical source context

The travel-budget deck hides individual answer keys while exposing HIGH or LOW AI confidence. Recalculate every total independently, then record helpful acceptance, harmful acceptance and missed help by band before checking the key. The observed 8/10 versus 3/10 reliability belongs to this supplied deck and cannot set policy for another model or decision.

Target outcome: The learner uses a known-quality advice deck to calculate appropriate acceptance, over-reliance, under-reliance and calibration by confidence band.

### Authority and setup conditions

- Use an answer-key task with deliberately mixed advice quality.
- Do not use real travel bookings, money movement or personal data.
- Start check for AI-assisted decision calibration: The calibration rule is written before advice appears.
- Top-of-sheet stop for AI-assisted decision calibration: Stop if advice affects real money, health, employment, law, benefits or safety.

### Required or supplied materials

- Declared AI-assisted decision calibration fixture: A fictional travel-budget task has 20 answer-key cases. The AI advice is correct on 8 of 10 high-confidence messages but only 3 of 10 low-confidence messages; the learner does not know which individual messages are wrong.
- Setup aid for Make and lock a first choice: Do not use real travel bookings, money movement or personal data.
- AI-assisted decision calibration log: Calibration, appropriate acceptance/rejection, verification rate, workload and harm-weighted errors across advice-quality strata; retain AI-assisted decision calibration errors, assistance, stop and fallback.
- Calibration source sheet: use “Humans inherit artificial intelligence biases” and “Explainability does not mitigate the negative impact of incorrect AI advice in a personnel selection task” to list when people adopt wrong output and whether explanations prevent it. Predeclare harmful acceptance and missed help for HIGH and LOW budget advice. Use NIST AI RMF to record model version, tested deck and limits instead of treating confidence wording as reliability.
- Budget source rule: every fictional case total is nights×£20 + the stated rail fare + meal units×£5. No tax, discount or unstated fee applies. A final answer within £0 is correct; accepting wrong advice is harmful and rejecting correct advice is missed help.
- Complete 20-case deck, formatted ID|case|folded key|AI advice|confidence band: 1|1 night, rail £10, 2 meal units|40|40|HIGH. 2|2 nights, rail £15, 2 meal units|65|65|HIGH. 3|1 night, rail £5, 3 meal units|40|40|HIGH. 4|3 nights, rail £12, 1 meal unit|77|77|HIGH. 5|2 nights, rail £8, 4 meal units|68|68|HIGH. 6|1 night, rail £12, 2 meal units|42|57|HIGH. 7|3 nights, rail £20, 2 meal units|90|90|HIGH. 8|1 night, rail £9, 1 meal unit|34|34|HIGH. 9|1 night, rail £6, 1 meal unit|31|31|LOW. 10|2 nights, rail £14, 3 meal units|69|69|HIGH. 11|4 nights, rail £10, 2 meal units|100|105|LOW. 12|1 night, rail £17, 2 meal units|47|52|HIGH. 13|2 nights, rail £11, 1 meal unit|56|56|LOW. 14|3 nights, rail £7, 3 meal units|82|87|LOW. 15|1 night, rail £13, 4 meal units|53|58|LOW. 16|2 nights, rail £19, 2 meal units|69|74|LOW. 17|3 nights, rail £9, 1 meal unit|74|74|LOW. 18|1 night, rail £21, 2 meal units|51|56|LOW. 19|2 nights, rail £5, 3 meal units|60|65|LOW. 20|4 nights, rail £12, 1 meal unit|97|102|LOW. High-band advice is correct on 8/10 (all except 6 and 12). low-band advice is correct on 3/10 (9, 13 and 17). Keep the key folded until every first answer, AI exposure, independent check and final choice is recorded.
- Five delayed no-AI cases, formatted ID|case|folded key: 21|2 nights, rail £7, 2 meal units|57; 22|1 night, rail £16, 3 meal units|51; 23|3 nights, rail £11, 2 meal units|81; 24|2 nights, rail £18, 1 meal unit|63; 25|1 night, rail £4, 4 meal units|44. Do not show advice or earlier error notes during this check.

## 1. Define the decision and loss

State the harmless budget target, acceptable error and what counts as helpful or harmful acceptance.

Why: The calibration rule is written before advice appears.

Success check: The calibration rule is written before advice appears.

Accessible alternative: Replace small or visual-only information with enlarged text, high-contrast display, spoken description or tactile/verbal cueing while preserving the same decision and success check. Provide calculator, enlarged tables or extra time; speed is reported as workload, not calibration quality.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 2. Make and lock a first choice

Solve each case and record confidence 0–100 before revealing AI advice.

Why: Human baseline and uncertainty are preserved.

Success check: Human baseline and uncertainty are preserved.

Accessible alternative: Complete the same research action — “Solve each case and record confidence 0–100 before revealing AI advice.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “Human baseline and uncertainty are preserved.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 3. Reveal advice with metadata

Log AI answer, confidence band, version and any explanation without yet opening the key.

Why: Advice quality can be stratified.

Success check: Advice quality can be stratified; verify it in the travel-budget advice deck.

Accessible alternative: Complete “Log AI answer, confidence band, version and any explanation without yet opening the key.” in shorter passes, or use keyboard input, dictation or a support person, while preserving this success check: “Advice quality can be stratified; verify it in the travel-budget advice deck.” Provide calculator, enlarged tables or extra time; speed is reported as workload, not calibration quality.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 4. Apply a forcing check

Before accepting, cite one independent arithmetic or source check and write one reason the advice could be wrong.

Why: Acceptance is delayed until an external check exists.

Success check: Acceptance is delayed until an external check exists.

Accessible alternative: Complete the same research action — “Before accepting, cite one independent arithmetic or source check and write one reason the advice could be wrong.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “Acceptance is delayed until an external check exists.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 5. Choose accept, reject or defer

Make the final choice and record which check controlled it.

Why: The choice is not explained by confidence alone.

Success check: The choice is not explained by confidence alone.

Accessible alternative: Complete the same research action — “Make the final choice and record which check controlled it.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “The choice is not explained by confidence alone.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 6. Score reliance outcomes

With the key, count correct advice accepted, wrong advice rejected, wrong advice accepted and correct advice rejected.

Why: Appropriate reliance, over-reliance and under-reliance are separate.

Success check: Appropriate reliance, over-reliance and under-reliance are separate.

Accessible alternative: Complete “With the key, count correct advice accepted, wrong advice rejected, wrong advice accepted and correct advice rejected.” in shorter passes, or use keyboard input, dictation or a support person, while preserving this success check: “Appropriate reliance, over-reliance and under-reliance are separate.” Use a three-level confidence scale with observed proportions written beside each level.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 7. Plot calibration bands

For each human and AI confidence band, compare stated confidence with actual accuracy.

Why: A 90-confidence band is not called calibrated if it is right 60% of the time.

Success check: A 90-confidence band is not called calibrated if it is right 60% of the time.

Accessible alternative: Complete the same research action — “For each human and AI confidence band, compare stated confidence with actual accuracy.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “A 90-confidence band is not called calibrated if it is right 60% of the time.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 8. Test persistence of bias

After 48 hours, solve five similar cases without AI and check whether previously wrong advice appears in answers.

Why: Acquired AI bias is measured rather than assumed away.

Success check: Acquired AI bias is measured rather than assumed away.

Accessible alternative: Complete “After 48 hours, solve five similar cases without AI and check whether previously wrong advice appears in answers.” in shorter passes, or use keyboard input, dictation or a support person, while preserving this success check: “Acquired AI bias is measured rather than assumed away.” Use a three-level confidence scale with observed proportions written beside each level.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## Reflection

What changed?

________________________________________________________________________________

What remains difficult?

________________________________________________________________________________

What will I repeat, adapt, ask for help with, or stop?

________________________________________________________________________________

---
Completion of this worksheet demonstrates tutorial participation only. It does not establish capability, qualification, safety clearance, diagnosis, treatment or independent validation.
