Use twelve questions to learn the confidence-calibration method
Confidence calibration asks whether answers given with a stated confidence are correct about that often. This lesson teaches the method on twelve objective two-choice questions: answer first, record confidence before feedback, preserve errors, group results by confidence and compare confidence with accuracy. Twelve items teach the calculation; they do not establish a stable trait or a reliable personal calibration score.
Tutorial curriculum · current public edition · Revision 7T · Full Tutorial Edition
One source of teaching truth
Full individual tutorial · TLU-PWR-121
Confidence calibration asks whether answers given with a stated confidence are correct about that often. This lesson teaches the method on twelve objective two-choice questions: answer first, record confidence before feedback, preserve errors, group results by confidence and compare confidence with accuracy. Twelve items teach the calculation; they do not establish a stable trait or a reliable personal calibration score.
Revision 7T · Full Tutorial Edition · complete governed practice tutorial
Use twelve questions to learn the confidence-calibration method
Confidence calibration asks whether answers given with a stated confidence are correct about that often. This lesson teaches the method on twelve objective two-choice questions: answer first, record confidence before feedback, preserve errors, group results by confidence and compare confidence with accuracy. Twelve items teach the calculation; they do not establish a stable trait or a reliable personal calibration score.
Authority and safety
Do not use this small benign set to rank people, infer intelligence, screen cognition or make school, work, clinical or eligibility decisions.
The supplied four-item example bins are deliberately small and teach arithmetic only. Report n every time; do not call a pattern stable until it has been repeated across a much larger set of resolved questions from the same topic.
Choose the answer, then record confidence, then lock both before opening the key. Never change either after feedback and never delete errors.
Exact materials
The twelve supplied questions with answer and confidence columns
The answer key, kept covered until all twelve answers and confidence ratings are locked
A calculation sheet or spreadsheet
For the delayed check, a fresh objective twelve-question set from the same topic with a trustworthy hidden key; do not reuse these questions
Set up in this order
Cover the answer key. Use the same confidence labels for the whole lesson: 50%, 60%, 70%, 80%, 90% or 100%.
For each question, choose A or B first. Immediately record how confident you are, then lock both entries before moving on.
After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong. Preserve every error.
Group questions that have the same confidence. For each group, record n, correct count and accuracy = correct divided by n.
For each group, calculate signed gap = confidence minus accuracy. A positive gap means confidence was higher than accuracy in that group; a negative gap means it was lower.
Write one feedback action limited to the same topic. Do not change future confidence from a four-item bin alone; keep collecting resolved questions.
After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic. Compare only confidence groups present in both records and report n beside every percentage.
Supplied fixture
Fill the final two columns before opening the answer key.
Item id
Question
Option a
Option b
Learner response before key
Learner confidence percent before key
Q1
What is 7 x 8?
54
56
[A/B]
[50/60/70/80/90/100]
Q2
Sum of a triangle's interior angles?
180 degrees
360 degrees
[A/B]
[50/60/70/80/90/100]
Q3
Which fraction equals 0.25?
1/3
1/4
[A/B]
[50/60/70/80/90/100]
Q4
Next: 2, 4, 8, 16?
24
32
[A/B]
[50/60/70/80/90/100]
Q5
What is 15% of 200?
30
35
[A/B]
[50/60/70/80/90/100]
Q6
Median of 2, 4, 9?
4
5
[A/B]
[50/60/70/80/90/100]
Q7
Metres in one kilometre?
100
1000
[A/B]
[50/60/70/80/90/100]
Q8
Is 3/5 larger than 2/3?
Yes
No
[A/B]
[50/60/70/80/90/100]
Q9
Area of a 3 by 4 rectangle?
12
14
[A/B]
[50/60/70/80/90/100]
Q10
Only even prime?
1
2
[A/B]
[50/60/70/80/90/100]
Q11
0.40 + 0.35?
0.75
0.85
[A/B]
[50/60/70/80/90/100]
Q12
Fair die probability of greater than 4?
1/2
1/3
[A/B]
[50/60/70/80/90/100]
Check your work only after you finish
Lock every response first. The checking page is separate so the answers are not exposed inside this exercise.
The method passes only when all twelve answers and confidence ratings were locked before the key, every item remains in the record, and every used confidence group shows n, correct count, accuracy and signed gap. A delayed comparison must use fresh questions from the same topic. This checks use of the method, not intelligence, wisdom or a stable confidence trait.
Worked example
Use the supplied fictional log for Q5–Q8: answers A, A, B and A, each recorded at 80% before the key.
The key shows Q5, Q6 and Q7 correct and Q8 wrong, so the group has n=4 and 3 correct.
Accuracy is 3 divided by 4 = 75%.
Signed gap is 80% minus 75% = +5 percentage points.
Because n=4, record this as a calculation example, not evidence that the learner is generally overconfident.
Keep future answers and confidence ratings in the same topic. Recalculate only after a larger resolved set, then test any feedback on fresh questions rather than editing old ratings.
Correct result: The fictional 80% group has 75% accuracy and a +5-point signed gap at n=4. The arithmetic is correct, but the group is too small for a stable personal conclusion.
Right and wrong
Right and wrong comparison
Moment
Right / safer
Wrong / riskier
Why
Recording sequence
Choose A or B, record confidence, then lock both before opening the key.
Look at the key, then decide how confident the answer felt.
A rating made after feedback cannot test calibration.
Confidence scale
Use only 50%, 60%, 70%, 80%, 90% or 100% throughout.
Switch between a six-point scale and only 60%, 80% and 100%.
Changing labels makes sessions and groups inconsistent.
Small groups
Write 75% accuracy, +5-point gap, n=4; calculation example only.
Four answers prove the learner is overconfident.
A tiny group is too noisy for a stable conclusion.
Wrong answers
Keep every wrong answer and its original confidence.
Delete misunderstood questions after seeing the key.
Selective deletion makes calibration look better than it was.
Delayed check
Use fresh questions from the same topic after at least seven days.
Repeat the same twelve remembered questions or compare unrelated topics.
Memory and topic difficulty would contaminate the comparison.
Common mistakes and fixes
Common mistakes and corrections
Mistake
Fix
Confidence is entered before choosing an answer or after seeing feedback.
Use the fixed order: answer, confidence, lock, key.
The page's fictional responses are mistaken for the learner's responses.
Complete the blank columns first; use the fictional log only for the worked calculation.
A confidence group is reported without n.
Write item count and correct count beside every accuracy and gap.
A four-item group drives an immediate confidence change.
Keep collecting fresh resolved questions from the same topic before changing the rule.
Old ratings are edited after feedback.
Preserve the old record and apply feedback only to fresh questions.
Evidence boundary
What the tutorial may—and may not—claim.
Supportable claim
Confidence can be scored against outcomes, but scalable calibration training has conflicting task-dependent evidence.
Measurement boundary
Declare question class, resolution rules, probability scale, calibration, resolution, base rate, scoring, sample size and missing outcomes; average confidence is not calibration.
Myth
Always know exactly how sure to be
Metric
Calibration and resolution over a predeclared set of resolved judgments
Boundary
Good calibration in one question class does not prove wisdom or transfer to another.
Negative and limiting findings
Practical-scoring feedback failed in two large experiments.
Automated feedback improved calibration in modified blackjack but not overall calibration in a more realistic baseball task.
Eric R. Stone; Jason Luu; Cory K. Costello; Annie H. Somerville · 2023 · PRIMARY_RESEARCH
Supports only the bounded empirical proposition in the declared configurations. Constrains generalisation, transfer, certainty or efficacy; it is not optional context.
Matthew Martin; David R. Mandel · 2024 · PRIMARY_RESEARCH
Supports only the bounded empirical proposition in the declared configurations. Constrains generalisation, transfer, certainty or efficacy; it is not optional context.
Zana Buçinca; Maja Barbara Malaya; Krzysztof Z. Gajos · 2021 · PRIMARY_RESEARCH
Supports only the bounded empirical proposition in the declared configurations. Constrains generalisation, transfer, certainty or efficacy; it is not optional context.
National Institute of Standards and Technology · 2023 · OFFICIAL_FRAMEWORK
Constrains generalisation, transfer, certainty or efficacy; it is not optional context. Supplies an official safety, access or governance boundary and is not empirical efficacy evidence.
Tutorial delivery controls
Learn, adapt, troubleshoot and resume
Estimated timeEstimated 19 min reading and worksheet pass
DifficultyIntroductory
EquipmentBasic stationery or digital tools
SpaceDesk / seated
Method qualityEstablished10 of 10 structural checks present. Automated method-readiness band; human editorial sign-off is separate.
Evidence contextG1; Detailed research depthScientific support is evaluated separately from teaching-method structure.
Editorial reviewPending manual sign-offNo human approval is claimed until reviewer, date and content hash are recorded.
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
01
Cover the answer key
Cover the answer key. Use the same confidence labels for the whole lesson: 50%, 60%, 70%, 80%, 90% or 100%.
Why this step exists
“Cover the answer key” carries the authored Confidence calibration method into a reviewable output without adding an unstated task, dose or claim.
Success check
Evidence of completion shows the learner followed this exact instruction without adding an unstated step: “Cover the answer key. Use the same confidence labels for the whole lesson: 50%, 60%, 70%, 80%, 90% or 100%.”
I’m stuck on this step
Reset: Re-read this authored instruction — “Cover the answer key. Use the same confidence labels for the whole lesson: 50%, 60%, 70%, 80%, 90% or 100%.” — and its success check, then attempt only this step.
Possible snag: Confidence is entered before choosing an answer or after seeing feedback.
Correction: Use the fixed order: answer, confidence, lock, key.
Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.
02
For each question, choose A or B first
For each question, choose A or B first. Immediately record how confident you are, then lock both entries before moving on.
Why this step exists
“For each question, choose A or B first” fixes the exact Confidence calibration configuration or decision before later observations are compared.
Success check
A legible entry directly completes “For each question, choose A or B first” and records the requested condition, decision or boundary without adding a new claim.
I’m stuck on this step
Reset: Re-read this authored instruction — “For each question, choose A or B first. Immediately record how confident you are, then lock both entries before moving on.” — and its success check, then attempt only this step.
Possible snag: The page's fictional responses are mistaken for the learner's responses.
Correction: Complete the blank columns first; use the fictional log only for the worked calculation.
Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.
03
After all twelve entries are locked, open the key and score every…
After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong. Preserve every error.
Why this step exists
“After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong” creates the auditable value, condition or statement needed for the tutorial’s later comparison and conclusion.
Success check
The work shows the value requested by “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong”, retains the supplied units or signs and leaves any unsupported value unknown.
I’m stuck on this step
Reset: Re-read this authored instruction — “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong. Preserve every error.” — and its success check, then attempt only this step.
Possible snag: The result from “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong. Preserve every error.” does not yet meet this declared check: The work shows the value requested by “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong”, retains the supplied units or signs and leaves any unsupported value unknown.
Correction: Return to the start of “After all twelve entries are locked, open the key and score every…”, reduce complexity or pace, and repeat only the part needed to satisfy: “The work shows the value requested by “After all twelve entries are locked, open the key and score every answer 1 for correct or 0 for wrong”, retains the supplied units or signs and leaves any unsupported value unknown.”
Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.
04
Group questions that have the same confidence
Group questions that have the same confidence. For each group, record n, correct count and accuracy = correct divided by n.
Why this step exists
“Group questions that have the same confidence” creates the auditable value, condition or statement needed for the tutorial’s later comparison and conclusion.
Success check
A legible entry directly completes “Group questions that have the same confidence” and records the requested condition, decision or boundary without adding a new claim.
I’m stuck on this step
Reset: Re-read this authored instruction — “Group questions that have the same confidence. For each group, record n, correct count and accuracy = correct divided by n.” — and its success check, then attempt only this step.
Possible snag: A confidence group is reported without n.
Correction: Write item count and correct count beside every accuracy and gap.
Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.
05
For each group, calculate signed gap = confidence minus accuracy
For each group, calculate signed gap = confidence minus accuracy. A positive gap means confidence was higher than accuracy in that group; a negative gap means it was lower.
Why this step exists
“For each group, calculate signed gap = confidence minus accuracy” creates the auditable value, condition or statement needed for the tutorial’s later comparison and conclusion.
Success check
The work shows the value requested by “For each group, calculate signed gap = confidence minus accuracy”, retains the supplied units or signs and leaves any unsupported value unknown.
I’m stuck on this step
Reset: Re-read this authored instruction — “For each group, calculate signed gap = confidence minus accuracy. A positive gap means confidence was higher than accuracy in that group; a negative gap means it was lower.” — and its success check, then attempt only this step.
Possible snag: The result from “For each group, calculate signed gap = confidence minus accuracy. A positive gap means confidence was higher than accuracy in that group; a negative gap means it was lower.” does not yet meet this declared check: The work shows the value requested by “For each group, calculate signed gap = confidence minus accuracy”, retains the supplied units or signs and leaves any unsupported value unknown.
Correction: Return to the start of “For each group, calculate signed gap = confidence minus accuracy”, reduce complexity or pace, and repeat only the part needed to satisfy: “The work shows the value requested by “For each group, calculate signed gap = confidence minus accuracy”, retains the supplied units or signs and leaves any unsupported value unknown.”
Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.
06
Write one feedback action limited to the same topic
Write one feedback action limited to the same topic. Do not change future confidence from a four-item bin alone; keep collecting resolved questions.
Why this step exists
“Write one feedback action limited to the same topic” creates the auditable value, condition or statement needed for the tutorial’s later comparison and conclusion.
Success check
A legible entry directly completes “Write one feedback action limited to the same topic” and records the requested condition, decision or boundary without adding a new claim.
I’m stuck on this step
Reset: Re-read this authored instruction — “Write one feedback action limited to the same topic. Do not change future confidence from a four-item bin alone; keep collecting resolved questions.” — and its success check, then attempt only this step.
Possible snag: A four-item group drives an immediate confidence change.
Correction: Keep collecting fresh resolved questions from the same topic before changing the rule.
Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.
07
After at least seven days, repeat the sequence on a fresh twelve-question…
After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic. Compare only confidence groups present in both records and report n beside every percentage.
Why this step exists
“After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic” isolates the named evidence distinction so the learner can interpret Confidence calibration without extending the claim.
Success check
The output directly answers “After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic”, names the relevant distinction and stays within the supplied record.
I’m stuck on this step
Reset: Re-read this authored instruction — “After at least seven days, repeat the sequence on a fresh twelve-question set from the same topic. Compare only confidence groups present in both records and report n beside every percentage.” — and its success check, then attempt only this step.
Possible snag: Old ratings are edited after feedback.
Correction: Preserve the old record and apply feedback only to fresh questions.
Stop / get help: A governed step-by-step lesson may be completed independently inside its stated limits. Stop if the declared configuration cannot be maintained, uncertainty becomes material, adverse effects appear, or qualified authority is required.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Recording sequence — A rating made after feedback cannot test calibration.
Correct / safer
Choose A or B, record confidence, then lock both before opening the key.
Wrong / riskier
Look at the key, then decide how confident the answer felt.
Confidence scale — Changing labels makes sessions and groups inconsistent.
Correct / safer
Use only 50%, 60%, 70%, 80%, 90% or 100% throughout.
Wrong / riskier
Switch between a six-point scale and only 60%, 80% and 100%.
Small groups — A tiny group is too noisy for a stable conclusion.
Correct / safer
Write 75% accuracy, +5-point gap, n=4; calculation example only.
Wrong / riskier
Four answers prove the learner is overconfident.
Wrong answers — Selective deletion makes calibration look better than it was.
Correct / safer
Keep every wrong answer and its original confidence.
Wrong / riskier
Delete misunderstood questions after seeing the key.
Delayed check — Memory and topic difficulty would contaminate the comparison.
Correct / safer
Use fresh questions from the same topic after at least seven days.
Wrong / riskier
Repeat the same twelve remembered questions or compare unrelated topics.
Method-structure checklist
10 of 10 structural checks present
✓ Ordered, Power-specific instructions — present
✓ Every activity has a success check — present
✓ Materials or supplied records are declared — present
✓ Measurement or assessment rule is present — present
✓ Tutorial-specific troubleshooting is present — present
✓ Stopping or escalation boundary is present — present
✓ Every activity has an adjacent alternative — present
✓ Correct-versus-incorrect comparison is present — present
✓ Evidence context is bound to the Power record — present
✓ Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.