Section 303 of 440

Complete canonical tutorial. This reader section contains the same teaching body as PWR-121 · Confidence calibration. Open the Power dossier.

PWR-121 · Metacognition & Epistemic Resilience

Use twelve questions to learn the confidence-calibration method

Confidence calibration asks whether answers given with a stated confidence are correct about that often. This lesson teaches the method on twelve objective two-choice questions: answer first, record confidence before feedback, preserve errors, group results by confidence and compare confidence with accuracy. Twelve items teach the calculation; they do not establish a stable trait or a reliable personal calibration score.

Tutorial curriculum · current public edition · Revision 7T · Full Tutorial Edition

Revision 7T · Full Tutorial Edition · complete governed practice tutorial

Use twelve questions to learn the confidence-calibration method

Confidence calibration asks whether answers given with a stated confidence are correct about that often. This lesson teaches the method on twelve objective two-choice questions: answer first, record confidence before feedback, preserve errors, group results by confidence and compare confidence with accuracy. Twelve items teach the calculation; they do not establish a stable trait or a reliable personal calibration score.

Authority and safety

  • Do not use this small benign set to rank people, infer intelligence, screen cognition or make school, work, clinical or eligibility decisions.
  • The supplied four-item example bins are deliberately small and teach arithmetic only. Report n every time; do not call a pattern stable until it has been repeated across a much larger set of resolved questions from the same topic.
  • Choose the answer, then record confidence, then lock both before opening the key. Never change either after feedback and never delete errors.

Exact materials

  • The twelve supplied questions with answer and confidence columns
  • The answer key, kept covered until all twelve answers and confidence ratings are locked
  • A calculation sheet or spreadsheet
  • For the delayed check, a fresh objective twelve-question set from the same topic with a trustworthy hidden key; do not reuse these questions

Set up in this order

  1. Cover the answer key. Use the same confidence labels for the whole lesson: 50%, 60%, 70%, 80%, 90% or 100%.
  2. For each question, choose A or B first. Immediately record how confident you are, then lock both entries before moving on.
  3. After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong. Preserve every error.
  4. Group questions that have the same confidence. For each group, record n, correct count and accuracy = correct divided by n.
  5. For each group, calculate signed gap = confidence minus accuracy. A positive gap means confidence was higher than accuracy in that group; a negative gap means it was lower.
  6. Write one feedback action limited to the same topic. Do not change future confidence from a four-item bin alone; keep collecting resolved questions.
  7. After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic. Compare only confidence groups present in both records and report n beside every percentage.

Supplied fixture

Fill the final two columns before opening the answer key.
Item idQuestionOption aOption bLearner response before keyLearner confidence percent before key
Q1What is 7 x 8?5456[A/B][50/60/70/80/90/100]
Q2Sum of a triangle's interior angles?180 degrees360 degrees[A/B][50/60/70/80/90/100]
Q3Which fraction equals 0.25?1/31/4[A/B][50/60/70/80/90/100]
Q4Next: 2, 4, 8, 16?2432[A/B][50/60/70/80/90/100]
Q5What is 15% of 200?3035[A/B][50/60/70/80/90/100]
Q6Median of 2, 4, 9?45[A/B][50/60/70/80/90/100]
Q7Metres in one kilometre?1001000[A/B][50/60/70/80/90/100]
Q8Is 3/5 larger than 2/3?YesNo[A/B][50/60/70/80/90/100]
Q9Area of a 3 by 4 rectangle?1214[A/B][50/60/70/80/90/100]
Q10Only even prime?12[A/B][50/60/70/80/90/100]
Q110.40 + 0.35?0.750.85[A/B][50/60/70/80/90/100]
Q12Fair die probability of greater than 4?1/21/3[A/B][50/60/70/80/90/100]

Check your work only after you finish

Lock every response first. The checking page is separate so the answers are not exposed inside this exercise.

Open the separate checking page →

Scoring rule

The method passes only when all twelve answers and confidence ratings were locked before the key, every item remains in the record, and every used confidence group shows n, correct count, accuracy and signed gap. A delayed comparison must use fresh questions from the same topic. This checks use of the method, not intelligence, wisdom or a stable confidence trait.

Worked example

  1. Use the supplied fictional log for Q5–Q8: answers A, A, B and A, each recorded at 80% before the key.
  2. The key shows Q5, Q6 and Q7 correct and Q8 wrong, so the group has n=4 and 3 correct.
  3. Accuracy is 3 divided by 4 = 75%.
  4. Signed gap is 80% minus 75% = +5 percentage points.
  5. Because n=4, record this as a calculation example, not evidence that the learner is generally overconfident.
  6. Keep future answers and confidence ratings in the same topic. Recalculate only after a larger resolved set, then test any feedback on fresh questions rather than editing old ratings.

Correct result: The fictional 80% group has 75% accuracy and a +5-point signed gap at n=4. The arithmetic is correct, but the group is too small for a stable personal conclusion.

Right and wrong

Right and wrong comparison
MomentRight / saferWrong / riskierWhy
Recording sequenceChoose A or B, record confidence, then lock both before opening the key.Look at the key, then decide how confident the answer felt.A rating made after feedback cannot test calibration.
Confidence scaleUse only 50%, 60%, 70%, 80%, 90% or 100% throughout.Switch between a six-point scale and only 60%, 80% and 100%.Changing labels makes sessions and groups inconsistent.
Small groupsWrite 75% accuracy, +5-point gap, n=4; calculation example only.Four answers prove the learner is overconfident.A tiny group is too noisy for a stable conclusion.
Wrong answersKeep every wrong answer and its original confidence.Delete misunderstood questions after seeing the key.Selective deletion makes calibration look better than it was.
Delayed checkUse fresh questions from the same topic after at least seven days.Repeat the same twelve remembered questions or compare unrelated topics.Memory and topic difficulty would contaminate the comparison.

Common mistakes and fixes

Common mistakes and corrections
MistakeFix
Confidence is entered before choosing an answer or after seeing feedback.Use the fixed order: answer, confidence, lock, key.
The page's fictional responses are mistaken for the learner's responses.Complete the blank columns first; use the fictional log only for the worked calculation.
A confidence group is reported without n.Write item count and correct count beside every accuracy and gap.
A four-item group drives an immediate confidence change.Keep collecting fresh resolved questions from the same topic before changing the rule.
Old ratings are edited after feedback.Preserve the old record and apply feedback only to fresh questions.

Evidence boundary

What the tutorial may—and may not—claim.

Supportable claim

Confidence can be scored against outcomes, but scalable calibration training has conflicting task-dependent evidence.

Measurement boundary

Declare question class, resolution rules, probability scale, calibration, resolution, base rate, scoring, sample size and missing outcomes; average confidence is not calibration.

Myth

Always know exactly how sure to be

Metric

Calibration and resolution over a predeclared set of resolved judgments

Boundary

Good calibration in one question class does not prove wisdom or transfer to another.

Negative and limiting findings

  • Practical-scoring feedback failed in two large experiments.
  • Automated feedback improved calibration in modified blackjack but not overall calibration in a more realistic baseball task.

Prohibited wording

  • always know what you do not know
  • eliminate overconfidence
  • calibration score proves wisdom
  • AI gives objective confidence

Direct source register

Evidence that supports—and limits—the route.

Automated calibration training for forecasters

Eric R. Stone; Jason Luu; Cory K. Costello; Annie H. Somerville · 2023 · PRIMARY_RESEARCH

Supports only the bounded empirical proposition in the declared configurations. Constrains generalisation, transfer, certainty or efficacy; it is not optional context.

Calibration Feedback With the Practical Scoring Rule Does Not Improve Calibration of Confidence

Matthew Martin; David R. Mandel · 2024 · PRIMARY_RESEARCH

Supports only the bounded empirical proposition in the declared configurations. Constrains generalisation, transfer, certainty or efficacy; it is not optional context.

To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making

Zana Buçinca; Maja Barbara Malaya; Krzysztof Z. Gajos · 2021 · PRIMARY_RESEARCH

Supports only the bounded empirical proposition in the declared configurations. Constrains generalisation, transfer, certainty or efficacy; it is not optional context.

Artificial Intelligence Risk Management Framework (AI RMF 1.0)

National Institute of Standards and Technology · 2023 · OFFICIAL_FRAMEWORK

Constrains generalisation, transfer, certainty or efficacy; it is not optional context. Supplies an official safety, access or governance boundary and is not empirical efficacy evidence.

Tutorial delivery controls

Learn, adapt, troubleshoot and resume

Estimated timeEstimated 19 min reading and worksheet pass
DifficultyIntroductory
EquipmentBasic stationery or digital tools
SpaceDesk / seated
Method qualityEstablished10 of 10 structural checks present. Automated method-readiness band; human editorial sign-off is separate.
Evidence contextG1; Detailed research depthScientific support is evaluated separately from teaching-method structure.
Editorial reviewPending manual sign-offNo human approval is claimed until reviewer, date and content hash are recorded.
Your tutorial progress0 of 7 steps complete
0 of 7 steps complete
Download learner worksheet

Progress is saved only in this browser on this device.

Step-by-step learner mode

Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.

01

Cover the answer key

Cover the answer key. Use the same confidence labels for the whole lesson: 50%, 60%, 70%, 80%, 90% or 100%.

Why this step exists

“Cover the answer key” carries the authored Confidence calibration method into a reviewable output without adding an unstated task, dose or claim.

Success check

Evidence of completion shows the learner followed this exact instruction without adding an unstated step: “Cover the answer key. Use the same confidence labels for the whole lesson: 50%, 60%, 70%, 80%, 90% or 100%.”

I’m stuck on this step

Reset: Re-read this authored instruction — “Cover the answer key. Use the same confidence labels for the whole lesson: 50%, 60%, 70%, 80%, 90% or 100%.” — and its success check, then attempt only this step.

  1. Possible snag: Confidence is entered before choosing an answer or after seeing feedback.

    Correction: Use the fixed order: answer, confidence, lock, key.

Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.

02

For each question, choose A or B first

For each question, choose A or B first. Immediately record how confident you are, then lock both entries before moving on.

Why this step exists

“For each question, choose A or B first” fixes the exact Confidence calibration configuration or decision before later observations are compared.

Success check

A legible entry directly completes “For each question, choose A or B first” and records the requested condition, decision or boundary without adding a new claim.

I’m stuck on this step

Reset: Re-read this authored instruction — “For each question, choose A or B first. Immediately record how confident you are, then lock both entries before moving on.” — and its success check, then attempt only this step.

  1. Possible snag: The page's fictional responses are mistaken for the learner's responses.

    Correction: Complete the blank columns first; use the fictional log only for the worked calculation.

Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.

03

After all twelve entries are locked, open the key and score every…

After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong. Preserve every error.

Why this step exists

“After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong” creates the auditable value, condition or statement needed for the tutorial’s later comparison and conclusion.

Success check

The work shows the value requested by “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong”, retains the supplied units or signs and leaves any unsupported value unknown.

I’m stuck on this step

Reset: Re-read this authored instruction — “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong. Preserve every error.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong. Preserve every error.” does not yet meet this declared check: The work shows the value requested by “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong”, retains the supplied units or signs and leaves any unsupported value unknown.

    Correction: Return to the start of “After all twelve entries are locked, open the key and score every…”, reduce complexity or pace, and repeat only the part needed to satisfy: “The work shows the value requested by “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong”, retains the supplied units or signs and leaves any unsupported value unknown.”

Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.

04

Group questions that have the same confidence

Group questions that have the same confidence. For each group, record n, correct count and accuracy = correct divided by n.

Why this step exists

“Group questions that have the same confidence” creates the auditable value, condition or statement needed for the tutorial’s later comparison and conclusion.

Success check

A legible entry directly completes “Group questions that have the same confidence” and records the requested condition, decision or boundary without adding a new claim.

I’m stuck on this step

Reset: Re-read this authored instruction — “Group questions that have the same confidence. For each group, record n, correct count and accuracy = correct divided by n.” — and its success check, then attempt only this step.

  1. Possible snag: A confidence group is reported without n.

    Correction: Write item count and correct count beside every accuracy and gap.

Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.

05

For each group, calculate signed gap = confidence minus accuracy

For each group, calculate signed gap = confidence minus accuracy. A positive gap means confidence was higher than accuracy in that group; a negative gap means it was lower.

Why this step exists

“For each group, calculate signed gap = confidence minus accuracy” creates the auditable value, condition or statement needed for the tutorial’s later comparison and conclusion.

Success check

The work shows the value requested by “For each group, calculate signed gap = confidence minus accuracy”, retains the supplied units or signs and leaves any unsupported value unknown.

I’m stuck on this step

Reset: Re-read this authored instruction — “For each group, calculate signed gap = confidence minus accuracy. A positive gap means confidence was higher than accuracy in that group; a negative gap means it was lower.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “For each group, calculate signed gap = confidence minus accuracy. A positive gap means confidence was higher than accuracy in that group; a negative gap means it was lower.” does not yet meet this declared check: The work shows the value requested by “For each group, calculate signed gap = confidence minus accuracy”, retains the supplied units or signs and leaves any unsupported value unknown.

    Correction: Return to the start of “For each group, calculate signed gap = confidence minus accuracy”, reduce complexity or pace, and repeat only the part needed to satisfy: “The work shows the value requested by “For each group, calculate signed gap = confidence minus accuracy”, retains the supplied units or signs and leaves any unsupported value unknown.”

Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.

06

Write one feedback action limited to the same topic

Write one feedback action limited to the same topic. Do not change future confidence from a four-item bin alone; keep collecting resolved questions.

Why this step exists

“Write one feedback action limited to the same topic” creates the auditable value, condition or statement needed for the tutorial’s later comparison and conclusion.

Success check

A legible entry directly completes “Write one feedback action limited to the same topic” and records the requested condition, decision or boundary without adding a new claim.

I’m stuck on this step

Reset: Re-read this authored instruction — “Write one feedback action limited to the same topic. Do not change future confidence from a four-item bin alone; keep collecting resolved questions.” — and its success check, then attempt only this step.

  1. Possible snag: A four-item group drives an immediate confidence change.

    Correction: Keep collecting fresh resolved questions from the same topic before changing the rule.

Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.

07

After at least seven days, repeat the sequence on a fresh twelve-question…

After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic. Compare only confidence groups present in both records and report n beside every percentage.

Why this step exists

“After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic” isolates the named evidence distinction so the learner can interpret Confidence calibration without extending the claim.

Success check

The output directly answers “After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic”, names the relevant distinction and stays within the supplied record.

I’m stuck on this step

Reset: Re-read this authored instruction — “After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic. Compare only confidence groups present in both records and report n beside every percentage.” — and its success check, then attempt only this step.

  1. Possible snag: Old ratings are edited after feedback.

    Correction: Preserve the old record and apply feedback only to fresh questions.

Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.

Correct versus incorrect execution

These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.

Recording sequence — A rating made after feedback cannot test calibration.
PWR-121 correct and incorrect comparison: Recording sequenceRecording sequence. Correct or safer: Choose A or B, record confidence, then lock both before opening the key.. Wrong or riskier: Look at the key, then decide how confident the answer felt.. Why: A rating made after feedback cannot test calibration.SITUATIONRecording sequenceCORRECT / SAFERChoose A or B, record confidence, then lock bothbefore opening the key.WRONG / RISKIERLook at the key, then decide how confident theanswer felt.YESNO
Correct / safer

Choose A or B, record confidence, then lock both before opening the key.

Wrong / riskier

Look at the key, then decide how confident the answer felt.

Confidence scale — Changing labels makes sessions and groups inconsistent.
PWR-121 correct and incorrect comparison: Confidence scaleConfidence scale. Correct or safer: Use only 50%, 60%, 70%, 80%, 90% or 100% throughout.. Wrong or riskier: Switch between a six-point scale and only 60%, 80% and 100%.. Why: Changing labels makes sessions and groups inconsistent.SITUATIONConfidence scaleCORRECT / SAFERUse only 50%, 60%, 70%, 80%, 90% or 100%throughout.WRONG / RISKIERSwitch between a six-point scale and only 60%,80% and 100%.YESNO
Correct / safer

Use only 50%, 60%, 70%, 80%, 90% or 100% throughout.

Wrong / riskier

Switch between a six-point scale and only 60%, 80% and 100%.

Small groups — A tiny group is too noisy for a stable conclusion.
PWR-121 correct and incorrect comparison: Small groupsSmall groups. Correct or safer: Write 75% accuracy, +5-point gap, n=4; calculation example only.. Wrong or riskier: Four answers prove the learner is overconfident.. Why: A tiny group is too noisy for a stable conclusion.SITUATIONSmall groupsCORRECT / SAFERWrite 75% accuracy, +5-point gap, n=4;calculation example only.WRONG / RISKIERFour answers prove the learner is overconfident.YESNO
Correct / safer

Write 75% accuracy, +5-point gap, n=4; calculation example only.

Wrong / riskier

Four answers prove the learner is overconfident.

Wrong answers — Selective deletion makes calibration look better than it was.
PWR-121 correct and incorrect comparison: Wrong answersWrong answers. Correct or safer: Keep every wrong answer and its original confidence.. Wrong or riskier: Delete misunderstood questions after seeing the key.. Why: Selective deletion makes calibration look better than it was.SITUATIONWrong answersCORRECT / SAFERKeep every wrong answer and its originalconfidence.WRONG / RISKIERDelete misunderstood questions after seeing thekey.YESNO
Correct / safer

Keep every wrong answer and its original confidence.

Wrong / riskier

Delete misunderstood questions after seeing the key.

Delayed check — Memory and topic difficulty would contaminate the comparison.
PWR-121 correct and incorrect comparison: Delayed checkDelayed check. Correct or safer: Use fresh questions from the same topic after at least seven days.. Wrong or riskier: Repeat the same twelve remembered questions or compare unrelated topics.. Why: Memory and topic difficulty would contaminate the comparison.SITUATIONDelayed checkCORRECT / SAFERUse fresh questions from the same topic after atleast seven days.WRONG / RISKIERRepeat the same twelve remembered questions orcompare unrelated topics.YESNO
Correct / safer

Use fresh questions from the same topic after at least seven days.

Wrong / riskier

Repeat the same twelve remembered questions or compare unrelated topics.

Method-structure checklist

10 of 10 structural checks present

  • Ordered, Power-specific instructions — present
  • Every activity has a success check — present
  • Materials or supplied records are declared — present
  • Measurement or assessment rule is present — present
  • Tutorial-specific troubleshooting is present — present
  • Stopping or escalation boundary is present — present
  • Every activity has an adjacent alternative — present
  • Correct-versus-incorrect comparison is present — present
  • Evidence context is bound to the Power record — present
  • Planning metadata is present — present

The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.

Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.