Complete bounded explanation
What this power means.
Capability to develop or express confidence calibration in a declared context without inheriting broader claims.
What the current evidence supports
Confidence can be scored against outcomes, but scalable calibration training has conflicting task-dependent evidence.
How to observe or measure it without overclaiming
Declare question class, resolution rules, probability scale, calibration, resolution, base rate, scoring, sample size and missing outcomes; average confidence is not calibration.
Always know exactly how sure to be
MetricCalibration and resolution over a predeclared set of resolved judgments
BoundaryGood calibration in one question class does not prove wisdom or transfer to another.
Negative and limiting findings
- Practical-scoring feedback failed in two large experiments.
- Automated feedback improved calibration in modified blackjack but not overall calibration in a more realistic baseball task.
Claims this evidence cannot support
- always know what you do not know
- eliminate overconfidence
- calibration score proves wisdom
- AI gives objective confidence