Section 300 of 440
Complete canonical tutorial. This reader section contains the same teaching body as PWR-118 · Expert judgement. Open the Power dossier.
PWR-118 · SUPERVISED full tutorial
Calibrate domain forecasts against resolved cases under expert supervision
Expert judgement is domain-specific performance in a feedback environment. A qualified domain expert selects representative resolved cases, protects sensitive information, defines the outcome and teaches probability forecasts. The learner commits a probability before feedback, then checks calibration—whether 70% forecasts occur about 70% of the time—and discrimination—whether higher probabilities went to events that happened. Credentials, confidence and eloquence are not substituted for resolved accuracy.
1 · Permission and limits
Know exactly what you may do
2 · Get ready
Gather what you need and check the starting conditions
What you need
- Qualified domain instructor, 20–40 representative resolved and de-identified cases and outcome key.
- Probability scale, rationale sheet, calculator or scoring tool and calibration chart.
- Declared domain, case-mix, conflict, AI/version and appeal record.
Before you start
- Provider verifies scope, confidentiality, case representativeness and legitimate use.
- Define the binary or categorical outcome and the resolution date before any forecast.
- Agree that every case receives a probability, not only memorable successes.
3 · The method
Follow these steps in order
- Lock domain and case mix
Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.
Why: Expertise can disappear outside its domain or on a selective case set.
Check: The case list matches the written inclusion rule.
- Forecast independently
For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.
Why: Independent commitment preserves the learner’s judgement for calibration.
Check: A timestamped probability exists for every case.
- Write diagnostic reasons
List the base rate, two case-specific cues and one reason the forecast could fail.
Why: Reasons expose whether confidence comes from relevant information or persuasive story.
Check: Each forecast has base rate, cues and counterargument.
- Reveal and score
After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.
Why: A proper score rewards probability accuracy across all outcomes.
Check: Every valid case contributes to the declared score.
- Compare with a simple benchmark
Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.
Why: Expert judgement adds value only if it can be separated from an appropriate baseline, not merely from guessing.
Check: The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.
- Check calibration and discrimination
Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.
Why: Calibration and discrimination are distinct qualities.
Check: The chart shows both reliability by bin and case separation.
- Repair and retest
The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.
Why: Specific feedback can improve a component without claiming generic expertise.
Check: The retest preserves case mix and documents one targeted repair.
4 · Worked example
See the whole method used once
Scenario
A qualified meteorological trainer gives a learner 20 archived, region-matched forecasts of measurable rain by the next day.
Walkthrough
- They define rain as at least the official recorded threshold at the named station and exclude severe-weather operations.
- Elena records a probability for every case before seeing outcome or model forecast.
- For one case she uses a 30% base rate, moisture cues and the counterargument that the front may stall, then forecasts 60%.
- Rain occurs on the 60% case, so she converts 60% to 0.60 and records (0.60−1)² = 0.16; she repeats that calculation for every resolved case.
- Her 20-case mean is 0.21 versus 0.24 for the 30% base-rate benchmark; in her 70–89% bin the mean forecast is 0.78 but only 0.60 of cases occur.
- The trainer identifies high-bin overconfidence, teaches a base-rate anchor and gives 20 fresh archived cases without calling the first difference proof of improvement.
Result
Elena can reproduce the Brier calculation, compare it on the same cases with the benchmark and identify a 0.18 high-bin calibration gap. Those descriptive results belong to this region, horizon and case mix and do not qualify her for live warnings or another domain.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
| Moment | Right / safer | Wrong / riskier | Why it matters |
|---|---|---|---|
| Case selection | Use a declared representative resolved sample. | Choose famous cases where the expert was right. | Selection bias can create false expertise. |
| Commitment | Forecast before outcome, peer or AI view. | Adjust the remembered forecast after resolution. | Hindsight destroys calibration evidence. |
| Metric | Score all probabilities against outcomes. | Judge confidence, credentials or eloquent rationale. | Persuasion can diverge from accuracy. |
| Transfer | Retest the same domain and label new domains separately. | Assume calibrated weather judgement transfers to finance or medicine. | Expert performance is domain- and feedback-specific. |
6 · Common mistakes
Spot the error and apply the correction
| Mistake | Fix |
|---|---|
| Only confident cases receive forecasts. | Require an entry for every sampled case. |
| Brier score is calculated after rounding probabilities to yes/no. | Use the original 0–1 probability for squared error. |
| Wide calibration bins hide overconfidence. | Predeclare bins and report counts in each. |
| AI advice contaminates the baseline. | Capture an independent forecast first and log any later model/version and update. |
7 · Practice
Turn the steps into a usable skill
First session
- Provider defines domain and resolution rule.
- Complete five familiarisation cases.
- Forecast a 20-case block independently.
- Score calibration and discrimination.
- Review one error class and set a matched retest.
Repeat plan
The provider schedules fresh resolved blocks until the sample is adequate and feedback remains representative. Live or operational cases are not used for unsupervised practice. Retest after delay and whenever the case mix, model or domain changes.
Progress when
- Calibration improves across enough provider-approved cases.
- Discrimination improves without cherry-picking or narrower case mix.
- The learner states scope, uncertainty, conflicts and escalation accurately.
Do not progress when
- Ground truth is ambiguous or outcomes are selectively missing.
- Practice cases leak into live consequential decisions.
- Improvement depends on seeing adviser or AI output before commitment.
8 · Check the result
Measure what changed
Provider-reviewed probabilistic accuracy on resolved domain cases.
How: For each resolved case calculate (p−o)² using forecast p on 0–1 and outcome o in {0,1}; average across cases for the Brier score. For each predeclared bin report mean p, occurred/total and their gap. Report mean p for occurred versus non-occurred cases, case count, base rate and missing outcomes.
Good result: A fresh matched block shows provider-defined improvement beyond sampling noise without narrowed case selection or hidden assistance.
This does not prove: It does not establish qualification, ethical authority, cross-domain expertise or safe live decision-making.
Self-check
- What exact domain and case mix are scored?
- Was every probability committed before feedback?
- Can you distinguish calibration from discrimination?
- Are conflicts and AI/version visible?
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
- Stop if a training answer could affect a live person, asset or safety outcome.
- Stop for confidentiality breach, unclear ground truth, conflict of interest or unlogged AI advice.
- Route high-stakes judgements through the accountable qualified system, including override and appeal.
Accessibility and adaptations
- Use plain-language cases, screen-reader-compatible probability controls and extended equal time.
- Allow a scribe who records without advising.
- Measure performance with declared supports rather than removing them.
10 · Evidence and limits
Why these instructions are here
- primary research
Geopolitical forecasting research identified task-specific drivers of prediction accuracy and calibration rather than credentials alone.
The psychology of intelligence analysis: Drivers of prediction accuracy in world politics - official guidance
NIST AI RMF treats human–AI performance as a governed, measured configuration with risks, not as automatic expert enhancement.
Artificial Intelligence Risk Management Framework (AI RMF 1.0)
Limits
- Resolved-case calibration belongs to the named domain and case mix.
- A low score does not itself grant professional authority.
- AI, credentials and persuasive reasons remain inputs to audit, not substitutes for outcomes.
Open the complete canonical research register
- Primary empirical supportLimiting / contraryThe psychology of intelligence analysis: Drivers of prediction accuracy in world politics
Barbara Mellers; Eric Stone; Pavel Atanasov; Nick Rohrbaugh; S. Emlen Metz; Lyle Ungar; Michael M. Bishop; Michael Horowitz; Ed Merkle; Philip Tetlock · 2015 · Primary research
- Primary empirical supportLimiting / contraryTo Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
Zana Buçinca; Maja Barbara Malaya; Krzysztof Z. Gajos · 2021 · Primary research
- Limiting / contraryDo People Listen to Cassandra? Persuasion and Accuracy in Geopolitical Forecasts
Yhonatan S. Shemesh; Avi Gamoran; David Leiser; Michael Gilead · 2025 · Primary research
- Primary empirical supportLimiting / contraryDebiasing Decisions
Carey K. Morewedge; Haewon Yoon; Irene Scopelliti; Carl W. Symborski; James H. Korris; Karim S. Kassam · 2015 · Primary research
- Primary empirical supportLimiting / contraryThe impact of AI errors in a human-in-the-loop process
Ujué Agudo; Karlos G. Liberal; Miren Arrese; Helena Matute · 2024 · Primary research
- Limiting / contraryOfficial boundary contextArtificial Intelligence Risk Management Framework (AI RMF 1.0)
National Institute of Standards and Technology · 2023 · Official framework
Read the complete evidence interpretation on the Power dossier.
Tutorial delivery controls
Learn, adapt, troubleshoot and resume
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
Lock domain and case mix
Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.
Expertise can disappear outside its domain or on a selective case set.
The case list matches the written inclusion rule.
I’m stuck on this step
Reset: Re-read this authored instruction — “Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.” — and its success check, then attempt only this step.
Possible snag: Only confident cases receive forecasts.
Correction: Require an entry for every sampled case.
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
Forecast independently
For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.
Independent commitment preserves the learner’s judgement for calibration.
A timestamped probability exists for every case.
I’m stuck on this step
Reset: Re-read this authored instruction — “For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.” — and its success check, then attempt only this step.
Possible snag: AI advice contaminates the baseline.
Correction: Capture an independent forecast first and log any later model/version and update.
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
Write diagnostic reasons
List the base rate, two case-specific cues and one reason the forecast could fail.
Reasons expose whether confidence comes from relevant information or persuasive story.
Each forecast has base rate, cues and counterargument.
I’m stuck on this step
Reset: Re-read this authored instruction — “List the base rate, two case-specific cues and one reason the forecast could fail.” — and its success check, then attempt only this step.
Possible snag: The result from “List the base rate, two case-specific cues and one reason the forecast could fail.” does not yet meet this declared check: Each forecast has base rate, cues and counterargument.
Correction: Return to the start of “Write diagnostic reasons”, reduce complexity or pace, and repeat only the part needed to satisfy: “Each forecast has base rate, cues and counterargument.”
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
Reveal and score
After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.
A proper score rewards probability accuracy across all outcomes.
Every valid case contributes to the declared score.
I’m stuck on this step
Reset: Re-read this authored instruction — “After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.” — and its success check, then attempt only this step.
Possible snag: Brier score is calculated after rounding probabilities to yes/no.
Correction: Use the original 0–1 probability for squared error.
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
Compare with a simple benchmark
Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.
Expert judgement adds value only if it can be separated from an appropriate baseline, not merely from guessing.
The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.
I’m stuck on this step
Reset: Re-read this authored instruction — “Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.” — and its success check, then attempt only this step.
Possible snag: The result from “Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.” does not yet meet this declared check: The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.
Correction: Return to the start of “Compare with a simple benchmark”, reduce complexity or pace, and repeat only the part needed to satisfy: “The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.”
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
Check calibration and discrimination
Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.
Calibration and discrimination are distinct qualities.
The chart shows both reliability by bin and case separation.
I’m stuck on this step
Reset: Re-read this authored instruction — “Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.” — and its success check, then attempt only this step.
Possible snag: Wide calibration bins hide overconfidence.
Correction: Predeclare bins and report counts in each.
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
Repair and retest
The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.
Specific feedback can improve a component without claiming generic expertise.
The retest preserves case mix and documents one targeted repair.
I’m stuck on this step
Reset: Re-read this authored instruction — “The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.” — and its success check, then attempt only this step.
Possible snag: The result from “The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.” does not yet meet this declared check: The retest preserves case mix and documents one targeted repair.
Correction: Return to the start of “Repair and retest”, reduce complexity or pace, and repeat only the part needed to satisfy: “The retest preserves case mix and documents one targeted repair.”
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Use a declared representative resolved sample.
Choose famous cases where the expert was right.
Forecast before outcome, peer or AI view.
Adjust the remembered forecast after resolution.
Score all probabilities against outcomes.
Judge confidence, credentials or eloquent rationale.
Retest the same domain and label new domains separately.
Assume calibrated weather judgement transfers to finance or medicine.
Method-structure checklist
10 of 10 structural checks present
- Ordered, Power-specific instructions — present
- Every activity has a success check — present
- Materials or supplied records are declared — present
- Measurement or assessment rule is present — present
- Tutorial-specific troubleshooting is present — present
- Stopping or escalation boundary is present — present
- Every activity has an adjacent alternative — present
- Correct-versus-incorrect comparison is present — present
- Evidence context is bound to the Power record — present
- Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.