PWR-121 · Metacognition & Epistemic Resilience

Confidence calibration

Calibration means matching confidence to outcomes—not simply becoming less confident.

Revision 7T · Full Tutorial Edition · Release manifest · Methodology · Corrections

One source of teaching truth

Canonical Power learning unit · TLU-PWR-121

Confidence calibration asks whether answers given with a stated confidence are correct about that often. This lesson teaches the method on twelve objective two-choice questions: answer first, record confidence before feedback, preserve errors, group results by confidence and compare confidence with accuracy. Twelve items teach the calculation; they do not establish a stable trait or a reliable personal calibration score.

Canonical Power page
PWR-121 · Confidence calibration
Full tutorial
Open full tutorial
Practical authority
Governed practice included inside the declared treatment
Current treatment
Self-guided lesson
Research depth
detailed · 4 bound sources
Risk framing
moderate
Capability self-practice
Permitted inside the governed tutorial limits
Pathway membership
G-CUR-012

Open My Power Path Inspect the canonical record

Read before using this page

Explanation is not permission.

This dossier explains the evidence and its limits. It is not a diagnosis, personal recommendation, assessment, clearance, performance promise or training programme. Actionable teaching appears only when the canonical Power record explicitly authorises it.

Complete bounded explanation

What this power means.

Capability to develop or express confidence calibration in a declared context without inheriting broader claims.

What the current evidence supports

Confidence can be scored against outcomes, but scalable calibration training has conflicting task-dependent evidence.

How to observe or measure it without overclaiming

Declare question class, resolution rules, probability scale, calibration, resolution, base rate, scoring, sample size and missing outcomes; average confidence is not calibration.

Myth

Always know exactly how sure to be

Metric

Calibration and resolution over a predeclared set of resolved judgments

Boundary

Good calibration in one question class does not prove wisdom or transfer to another.

Negative and limiting findings

  • Practical-scoring feedback failed in two large experiments.
  • Automated feedback improved calibration in modified blackjack but not overall calibration in a more realistic baseball task.

Claims this evidence cannot support

  • always know what you do not know
  • eliminate overconfidence
  • calibration score proves wisdom
  • AI gives objective confidence

Evidence guide frame · non-prescriptive

A guide frame—not a training protocol.

Explain confidence calibration against resolved outcomes without turning uncertainty into a personal score or universal wisdom claim.

What this frame may discuss

  • Evidence and measurement boundaries for confidence calibration
  • Task specificity, access configurations and negative findings
  • Limits on transfer, efficacy and authority

Conceptual observations

  • Conceptually distinguish average confidence, calibration and resolution.
  • Treat the question class, base rate, outcome rules and missing outcomes as part of the result.

Reflection prompts

  • What would count as a resolved outcome?
  • Could calibration in one question class transfer poorly to another?

What it cannot establish

  • Good calibration in one question class does not prove wisdom or transfer to another.
  • It cannot establish: always know what you do not know.
  • It cannot establish: eliminate overconfidence.
  • It cannot establish: calibration score proves wisdom.
  • It cannot establish: AI gives objective confidence.

Accessibility alternatives

  • Fictional benign judgments are sufficient; no submission or computation is needed.
  • Words, percentages, frequencies and visual scales are alternative representations.
  • No confidence diary, tracking schedule or score is requested.

Non-participation remains valid: Yes.

Stop and escalation boundaries

  • Stop if the frame becomes repeated self-scoring, a certainty target or an optimisation loop.
  • Do not use confidence labels as competence, credibility or clinical judgments about a person.
  • Consequential judgments require domain evidence, independent checks and accountable human review.

Prohibited uses

  • Calibration exercise, score, rank, target, diary or progression.
  • Clinical, employment, credibility or competence assessment.
  • Claims of universal wisdom or reliable transfer.
No collection or scoring: this page requests no answers, stores nothing and produces no personal or Titan result.

Open the complete governed guide-frame library entry

Individual tutorial · PRACTICAL LESSON

How to learn this Power now.

Confidence calibration asks whether answers given with a stated confidence are correct about that often. This lesson teaches the method on twelve objective two-choice questions: answer first, record confidence before feedback, preserve errors, group results by confidence and compare confidence with accuracy. Twelve items teach the calculation; they do not establish a stable trait or a reliable personal calibration score.

  1. OrientRead the definition, evidence, measurement boundary and prohibited claims on this canonical page.
  2. BaselineOpen the governed lesson and record its declared baseline before practising.
  3. PractiseFollow the ordered lesson exactly; do not add dose, intensity or claims.
  4. RetestUse the lesson's delayed retention and predeclared transfer checks as separate records.
  5. Maintain or retireKeep only what remains useful inside the evidence and safety boundary.

Authority boundary: A governed step-by-step lesson may be completed independently inside its stated limits. The method is Power-specific; its actionability follows this treatment.

Related Powers: PWR-122 · PWR-123 · PWR-124 · PWR-125

Direct evidence register

Sources that support—and limit—the claim.

Citation roles are explicit. A source may support existence or trainability while simultaneously limiting transfer, certainty, generalisation or safety.

  1. Primary empirical supportLimiting / contrary
    Automated calibration training for forecasters

    Eric R. Stone; Jason Luu; Cory K. Costello; Annie H. Somerville · 2023 · Primary research

  2. Primary empirical supportLimiting / contrary
    Calibration Feedback With the Practical Scoring Rule Does Not Improve Calibration of Confidence

    Matthew Martin; David R. Mandel · 2024 · Primary research

  3. Primary empirical supportLimiting / contrary
    To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making

    Zana Buçinca; Maja Barbara Malaya; Krzysztof Z. Gajos · 2021 · Primary research

  4. Limiting / contraryOfficial boundary context
    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology · 2023 · Official authority

Current teaching and evidence boundary

The next gate remains visible.

No external confirmation is required for the current bounded explainer and non-prescriptive guide-frame permissions; any stronger efficacy, generalisation, protocol, promotion or independent-validation claim requires a new adjudication.

A separately governed practical lesson is available at /tutorials/pwr-121/. Its safe configuration, prerequisites, ordered practice, measurement, stopping, accessibility, retention, transfer and review rules remain controlling.