Section 355 of 440
Complete canonical tutorial. This reader section contains the same teaching body as PWR-173 · Trust calibration. Open the Power dossier.
PWR-173 · SELF GUIDED full tutorial
Calibrate reliance on a fallible adviser by testing known reliability conditions
Use the weather-adviser deck to find out whether your reliance changes when advice quality falls from clear maps to blurry or reversed-key maps. Independent answers and the folded key expose helpful acceptance, harmful acceptance and needless rejection by condition. Those numbers belong to this fixed adviser version, not to people or future systems.
1 · Permission and limits
Know exactly what you may do
2 · Get ready
Gather what you need and check the starting conditions
What you need
- Declared Trust calibration fixture: A fictional weather adviser answers 24 map questions: it is correct on 6 of 8 clear-map items, 4 of 8 blurry-map items and 2 of 8 reversed-key items. An answer key permits exact scoring.
- Setup aid for Make an independent first answer: Predeclare the three reliability conditions and do not change them after seeing results.
- Trust calibration log: Appropriate reliance, over-reliance and under-reliance across known reliability conditions; retain Trust calibration errors, assistance, stop and fallback.
- Reliance evidence sheet: from “Calibrating Reliance on Automated Advice: Transparency and Trust Calibration Feedback”, extract the advice conditions and calibration feedback; from “Adaptive trust calibration for human-AI collaboration”, extract how changing reliability is represented. Compare those manipulations with the clear, blurry and reversed-key map strata. The worker-monitoring source supplies a privacy boundary, not permission to score people.
- Weather-adviser row format is ID|source-map prompt|folded correct shelter|adviser answer. Choose the shelter outside the marked weather before revealing advice.
- Clear rows: C1|rain band covers east shelter B; west shelter A is outside it|A|A; C2|rain band covers west shelter A; east shelter B is outside it|B|B
- Clear rows: C3|storm track crosses north shelter A; south shelter B is clear|B|B; C4|storm track crosses south shelter B; north shelter A is clear|A|A
- Clear rows: C5|flood mark reaches A only|B|B; C6|snow icon sits on B only|A|A
- Clear rows: C7|hail icon sits on A only|B|A; C8|fog icon sits on B only|A|B
- Blurry rows: B1|faint band boundary still crosses A; B is outside|B|B; B2|faded contour crosses B; A is outside|A|A
- Blurry rows: B3|faded snow icon is centred on A|B|B; B4|faded fog icon is centred on B|A|A
- Blurry rows: B5|faint rain mark is centred on A|B|A; B6|faint ice mark is centred on B|A|B
- Blurry rows: B7|faint hail mark is centred on A|B|A; B8|faint band is centred on B|A|B
- Reversed-key rows: R1|legend filled=clear/open=rain; A filled, B open|A|A; R2|legend filled=clear/open=rain; A open, B filled|B|B
- Reversed-key rows: R3|legend filled=clear/open=rain; A filled, B open|A|B; R4|legend filled=clear/open=rain; A open, B filled|B|A
- Reversed-key rows: R5|legend striped=clear/solid=rain; A striped, B solid|A|B; R6|legend striped=clear/solid=rain; A solid, B striped|B|A
- Reversed-key rows: R7|legend square=clear/triangle=rain; A square, B triangle|A|B; R8|legend square=clear/triangle=rain; A triangle, B square|B|A
- Baseline adviser accuracy is 6/8 on clear maps, 4/8 on blurry maps and 2/8 on reversed-key maps.
- Delayed rows: DC1|rain mark on B only|A|A; DC2|snow mark on A only|B|B
- Delayed rows: DC3|fog mark on B only|A|A; DC4|hail mark on A only|B|A
- Delayed rows: DB1|faint band on A only|B|B; DB2|faint band on B only|A|A
- Delayed rows: DB3|faint ice on A only|B|A; DB4|faint fog on B only|A|B
- Delayed rows: DR1|legend ring=clear/dot=rain; A ring, B dot|A|A; DR2|legend ring=clear/dot=rain; A dot, B ring|B|A
- Delayed rows: DR3|legend bar=clear/cross=rain; A bar, B cross|A|B; DR4|legend bar=clear/cross=rain; A cross, B bar|B|A
- Delayed advice accuracy is 3/4 clear, 2/4 blurry and 1/4 reversed-key. Keep every key covered until the independent answer, confidence, advice exposure and final choice are locked.
Before you start
- Use a fictional adviser and harmless answer-key task.
- Predeclare the three reliability conditions and do not change them after seeing results.
- Start check for Trust calibration: The rule names acceptance, rejection and the answer key.
- Top-of-sheet stop for Trust calibration: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
3 · The method
Follow these steps in order
- Define appropriate reliance
Write that advice should be accepted only when it improves the answer under the declared condition; expressed trust is not the target.
Why: The rule names acceptance, rejection and the answer key.
Check: The rule names acceptance, rejection and the answer key.
- Make an independent first answer
Answer each map question and rate confidence 0–100 before seeing the adviser.
Why: Every row has a pre-advice answer and confidence.
Check: Every row has a pre-advice answer and confidence.
- Reveal advice and condition
Show the adviser’s answer plus clear, blurry or reversed-key condition. Do not reveal correctness yet.
Why: Advice exposure and condition are logged without editing the first answer.
Check: Advice exposure and condition are logged without editing the first answer.
- Decide and justify
Keep or change the answer, citing map evidence and condition rather than liking the adviser.
Why: The final answer has a one-line evidence reason.
Check: The final answer has a one-line evidence reason.
- Score four outcomes
Use the key to count helpful acceptance, harmful acceptance, justified rejection and missed helpful advice.
Why: Over-reliance and under-reliance are separate counts.
Check: Over-reliance and under-reliance are separate counts.
- Compare reliability strata
Calculate adviser and team accuracy within each of the three conditions.
Why: Results are not averaged across conditions that have different reliability.
Check: Results are not averaged across conditions that have different reliability.
- Set a reliance rule
Choose a rule supported by the results, such as verify all reversed-key advice and inspect low-confidence disagreements on clear maps.
Why: The rule names the condition and verification action.
Check: The rule names the condition and verification action.
- Test without feedback
After 48 hours use a new 12-item set with the same strata and no item-by-item correction.
Why: Calibration persists on new items or is reported as feedback-dependent.
Check: Calibration persists on new items or is reported as feedback-dependent.
4 · Worked example
See the whole method used once
Scenario
A fictional weather adviser answers 24 map questions: it is correct on 6 of 8 clear-map items, 4 of 8 blurry-map items and 2 of 8 reversed-key items. An answer key permits exact scoring.
Walkthrough
- Answer the first eight clear-map questions without advice and lock confidence ratings.
- Reveal the fictional adviser’s answer, then keep or change each response with one map-based reason.
- Open the key and count one helpful acceptance, one harmful acceptance, two justified rejections and one missed helpful answer among disagreements.
- Repeat the count for blurry and reversed-key maps instead of pooling all 24.
- Adopt the rule “verify every reversed-key answer; inspect clear-map disagreements when my confidence is below 60.”
- Two days later, use 12 new questions and report whether harmful acceptance falls without missed-helpful advice rising.
Result
The log shows whether reliance matches known adviser reliability in this fixture. It does not establish trustworthiness of people, other systems or future versions.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
| Moment | Right / safer | Wrong / riskier | Why it matters |
|---|---|---|---|
| Independent judgement | Answer before seeing advice when completing the weather-adviser reliance deck. | Let the adviser’s wording become the starting answer. | Early-advice bias hides human-only performance and makes over-reliance hard to see. |
| Reliability by condition | Calculate accuracy separately for clear, blurry and reversed-key maps. | Use one overall “75% reliable” label. | Average reliability can hide a known failure condition. |
| Advice acceptance | Accept when verified evidence improves the answer. | Follow high-confidence advice automatically during the weather-adviser reliance deck. | Confidence is not correctness; this distorts the weather-adviser reliance deck. |
| Calibration result | Report harmful acceptance and missed helpful advice separately. | Call frequent acceptance “high trust” and therefore success. | Good calibration can require both acceptance and rejection. |
6 · Common mistakes
Spot the error and apply the correction
| Mistake | Fix |
|---|---|
| Let the adviser’s wording become the starting answer. | Lock the first answer and confidence before reveal. |
| Use one overall “75% reliable” label. | Keep one row and denominator per condition. |
| Follow high-confidence advice automatically during the weather-adviser reliance deck. | Cite the map feature or key rule before changing. |
| Call frequent acceptance “high trust” and therefore success. | Use the four-outcome table within the weather-adviser reliance deck. |
7 · Practice
Turn the steps into a usable skill
First session
- Weather-adviser reliance run: A fictional weather adviser answers 24 map questions: it is correct on 6 of 8 clear-map items, 4 of 8 blurry-map items and 2 of 8 reversed-key items. An answer key permits exact scoring.
- Independent map answer: Define appropriate reliance: Write that advice should be accepted only when it improves the answer under the declared condition; expressed trust is not the target.
- Condition-specific acceptance pass: Decide and justify, then Score four outcomes.
- Advice-bias correction: if “Let the adviser’s wording become the starting answer.” appears, apply “Lock the first answer and confidence before reveal.”
- Twelve-card delayed check: Test without feedback: After 48 hours use a new 12-item set with the same strata and no item-by-item correction.
Repeat plan
Run one 24-item calibrated deck, then a 12-item delayed deck after 48 hours. Repeat monthly with a changed reliability distribution; rewrite the rule whenever version or condition changes rather than carrying old trust forward.
Progress when
- The rule names acceptance, rejection and the answer key.
- Every row has a pre-advice answer and confidence.
- Advice exposure and condition are logged without editing the first answer.
- In the delayed deck, harmful acceptance is no worse than the independent baseline, helpful advice is sometimes used, and every reliance rule names a tested condition and version.
Do not progress when
- Do not continue while this error remains: Let the adviser’s wording become the starting answer.
- Pause until this correction works: Keep one row and denominator per condition.
- This Trust calibration stop ends the block: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
8 · Check the result
Measure what changed
Appropriate reliance, over-reliance and under-reliance across known reliability conditions
How: Baseline fixture: A fictional weather adviser answers 24 map questions: it is correct on 6 of 8 clear-map items, 4 of 8 blurry-map items and 2 of 8 reversed-key items. An answer key permits exact scoring. Enter success checks from “Define appropriate reliance” and “Score four outcomes”. If “Let the adviser’s wording become the starting answer.” occurs, apply its named fix; then score Appropriate reliance, over-reliance and under-reliance across known reliability conditions on the unused “Test without feedback” item.
Good result: In the delayed deck, harmful acceptance is no worse than the independent baseline, helpful advice is sometimes used, and every reliance rule names a tested condition and version.
This does not prove: Boundary for Trust calibration: “Appropriate reliance, over-reliance and under-reliance across known reliability conditions” describes only A fictional weather adviser answers 24 map questions: it is correct on 6 of 8 clear-map items, 4 of 8 blurry-map items and 2 of 8 reversed-key items. An answer key permits exact scoring. It cannot establish “Install a perfect trust detector”.
Self-check
- Without the example, demonstrate: The rule names acceptance, rejection and the answer key.
- Find the fault in this attempt: “Let the adviser’s wording become the starting answer.” Apply “Lock the first answer and confidence before reveal.”; what changes?
- What evidence in the completed record shows that this is wrong: “Use one overall “75% reliable” label.”?
- Trust calibration stop decision: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
- Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
- Do not use a live person’s trust as the answer key or manipulate a partner without consent.
- Re-baseline after any adviser version, data source or task change; old calibration does not transfer automatically.
Accessibility and adaptations
- Use large-print maps or text-only logic items with identical reliability strata and answer keys.
- Permit calculator use for rates; calibration is judged by decisions and denominators, not arithmetic fluency.
10 · Evidence and limits
Why these instructions are here
- primary research
Registered support for Trust calibration: “Calibrating Reliance on Automated Advice: Transparency and Trust Calibration Feedback”. It bears on Appropriate reliance, over-reliance and under-reliance across known reliability conditions inside the Trust calibration fixture. It does not validate “Install a perfect trust detector”.
Calibrating Reliance on Automated Advice: Transparency and Trust Calibration Feedback - primary research
Constraint for Trust calibration, drawn from “Adaptive trust calibration for human-AI collaboration”: Explicit trust-calibration feedback was null on key outcomes.
Adaptive trust calibration for human-AI collaboration - official guidance
ICO worker-monitoring limit — Trust calibration: rights and data protection apply to “Define appropriate reliance”. Fixture type for “Make an independent first answer”: fictional or consented. Covert use of Appropriate reliance, over-reliance and under-reliance across known reliability conditions is outside scope.
Employment practices and data protection: monitoring workers
Limits
- Trust calibration boundary: interpret “Appropriate reliance, over-reliance and under-reliance across known reliability conditions” only for A fictional weather adviser answers 24 map questions: it is correct on 6 of 8 clear-map items, 4 of 8 blurry-map items and 2 of 8 reversed-key items. An answer key permits exact scoring.
- A successful result does not establish “Install a perfect trust detector”.
- Trust calibration limiting finding: Explicit trust-calibration feedback was null on key outcomes.
- Trust calibration limiting finding: No device-off retention or cross-domain transfer was shown.
- Scope remains Trust calibration: A fictional weather adviser answers 24 map questions: it is correct on 6 of 8 clear-map items, 4 of 8 blurry-map items and 2 of 8 reversed-key items. An answer key permits exact scoring. Recheck the comparator, support and “Appropriate reliance, over-reliance and under-reliance across known reliability conditions” after any configuration change.
Open the complete canonical research register
- Primary empirical supportLimiting / contraryCalibrating Reliance on Automated Advice: Transparency and Trust Calibration Feedback
Monica Tatasciore; Shayne Loft · 2025 · Primary research
- Primary empirical supportLimiting / contraryAdaptive trust calibration for human-AI collaboration
Kazuo Okamura; Seiji Yamada · 2020 · Primary research
- Limiting / contraryOfficial boundary contextEmployment practices and data protection: monitoring workers
Information Commissioner’s Office · 2023 · Official guidance
Read the complete evidence interpretation on the Power dossier.
Tutorial delivery controls
Learn, adapt, troubleshoot and resume
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
Define appropriate reliance
Write that advice should be accepted only when it improves the answer under the declared condition; expressed trust is not the target.
The rule names acceptance, rejection and the answer key.
The rule names acceptance, rejection and the answer key.
I’m stuck on this step
Reset: Re-read this authored instruction — “Write that advice should be accepted only when it improves the answer under the declared condition; expressed trust is not the target.” — and its success check, then attempt only this step.
Possible snag: Call frequent acceptance “high trust” and therefore success.
Correction: Use the four-outcome table within the weather-adviser reliance deck.
Stop / get help: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
Make an independent first answer
Answer each map question and rate confidence 0–100 before seeing the adviser.
Every row has a pre-advice answer and confidence.
Every row has a pre-advice answer and confidence.
I’m stuck on this step
Reset: Re-read this authored instruction — “Answer each map question and rate confidence 0–100 before seeing the adviser.” — and its success check, then attempt only this step.
Possible snag: Let the adviser’s wording become the starting answer.
Correction: Lock the first answer and confidence before reveal.
Stop / get help: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
Reveal advice and condition
Show the adviser’s answer plus clear, blurry or reversed-key condition. Do not reveal correctness yet.
Advice exposure and condition are logged without editing the first answer.
Advice exposure and condition are logged without editing the first answer.
I’m stuck on this step
Reset: Re-read this authored instruction — “Show the adviser’s answer plus clear, blurry or reversed-key condition. Do not reveal correctness yet.” — and its success check, then attempt only this step.
Possible snag: The result from “Show the adviser’s answer plus clear, blurry or reversed-key condition. Do not reveal correctness yet.” does not yet meet this declared check: Advice exposure and condition are logged without editing the first answer.
Correction: Return to the start of “Reveal advice and condition”, reduce complexity or pace, and repeat only the part needed to satisfy: “Advice exposure and condition are logged without editing the first answer.”
Stop / get help: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
Decide and justify
Keep or change the answer, citing map evidence and condition rather than liking the adviser.
The final answer has a one-line evidence reason.
The final answer has a one-line evidence reason.
I’m stuck on this step
Reset: Re-read this authored instruction — “Keep or change the answer, citing map evidence and condition rather than liking the adviser.” — and its success check, then attempt only this step.
Possible snag: Use one overall “75% reliable” label.
Correction: Keep one row and denominator per condition.
Stop / get help: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
Score four outcomes
Use the key to count helpful acceptance, harmful acceptance, justified rejection and missed helpful advice.
Over-reliance and under-reliance are separate counts.
Over-reliance and under-reliance are separate counts.
I’m stuck on this step
Reset: Re-read this authored instruction — “Use the key to count helpful acceptance, harmful acceptance, justified rejection and missed helpful advice.” — and its success check, then attempt only this step.
Possible snag: The result from “Use the key to count helpful acceptance, harmful acceptance, justified rejection and missed helpful advice.” does not yet meet this declared check: Over-reliance and under-reliance are separate counts.
Correction: Return to the start of “Score four outcomes”, reduce complexity or pace, and repeat only the part needed to satisfy: “Over-reliance and under-reliance are separate counts.”
Stop / get help: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
Compare reliability strata
Calculate adviser and team accuracy within each of the three conditions.
Results are not averaged across conditions that have different reliability.
Results are not averaged across conditions that have different reliability.
I’m stuck on this step
Reset: Re-read this authored instruction — “Calculate adviser and team accuracy within each of the three conditions.” — and its success check, then attempt only this step.
Possible snag: The result from “Calculate adviser and team accuracy within each of the three conditions.” does not yet meet this declared check: Results are not averaged across conditions that have different reliability.
Correction: Return to the start of “Compare reliability strata”, reduce complexity or pace, and repeat only the part needed to satisfy: “Results are not averaged across conditions that have different reliability.”
Stop / get help: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
Set a reliance rule
Choose a rule supported by the results, such as verify all reversed-key advice and inspect low-confidence disagreements on clear maps.
The rule names the condition and verification action.
The rule names the condition and verification action.
I’m stuck on this step
Reset: Re-read this authored instruction — “Choose a rule supported by the results, such as verify all reversed-key advice and inspect low-confidence disagreements on clear maps.” — and its success check, then attempt only this step.
Possible snag: Follow high-confidence advice automatically during the weather-adviser reliance deck.
Correction: Cite the map feature or key rule before changing.
Stop / get help: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
Test without feedback
After 48 hours use a new 12-item set with the same strata and no item-by-item correction.
Calibration persists on new items or is reported as feedback-dependent.
Calibration persists on new items or is reported as feedback-dependent.
I’m stuck on this step
Reset: Re-read this authored instruction — “After 48 hours use a new 12-item set with the same strata and no item-by-item correction.” — and its success check, then attempt only this step.
Possible snag: The result from “After 48 hours use a new 12-item set with the same strata and no item-by-item correction.” does not yet meet this declared check: Calibration persists on new items or is reported as feedback-dependent.
Correction: Return to the start of “Test without feedback”, reduce complexity or pace, and repeat only the part needed to satisfy: “Calibration persists on new items or is reported as feedback-dependent.”
Stop / get help: Stop if the exercise is proposed for hiring, health, legal, financial or safety decisions.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Answer before seeing advice when completing the weather-adviser reliance deck.
Let the adviser’s wording become the starting answer.
Calculate accuracy separately for clear, blurry and reversed-key maps.
Use one overall “75% reliable” label.
Accept when verified evidence improves the answer.
Follow high-confidence advice automatically during the weather-adviser reliance deck.
Report harmful acceptance and missed helpful advice separately.
Call frequent acceptance “high trust” and therefore success.
Method-structure checklist
10 of 10 structural checks present
- Ordered, Power-specific instructions — present
- Every activity has a success check — present
- Materials or supplied records are declared — present
- Measurement or assessment rule is present — present
- Tutorial-specific troubleshooting is present — present
- Stopping or escalation boundary is present — present
- Every activity has an adjacent alternative — present
- Correct-versus-incorrect comparison is present — present
- Evidence context is bound to the Power record — present
- Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.