Revision 7T · Full Tutorial Edition · Updated 1 September 2026
PWR-118 · SUPERVISED full tutorial
Calibrate domain forecasts against resolved cases under expert supervision
Expert judgement is domain-specific performance in a feedback environment. A qualified domain expert selects representative resolved cases, protects sensitive information, defines the outcome and teaches probability forecasts. The learner commits a probability before feedback, then checks calibration—whether 70% forecasts occur about 70% of the time—and discrimination—whether higher probabilities went to events that happened. Credentials, confidence and eloquence are not substituted for resolved accuracy.
What you will produceUnder a qualified provider, the learner completes at least 20 resolved domain cases and produces a Brier score (mean squared probability error), calibration bands, discrimination and an error review for the declared case mix.
Method7 numbered Power-specific steps
Practice authorityFull method with qualified supervision where stated
Full step-by-step individual tutorial · TLU-PWR-118
Under a qualified provider, the learner completes at least 20 resolved domain cases and produces a Brier score (mean squared probability error), calibration bands, discrimination and an error review for the declared case mix.
State independent probability judgements, rationales, uncertainty and conflicts.
Use de-identified training cases selected by the provider.
Qualified help is required for
Defining domain scope and ground truth, selecting cases, protecting data, teaching standards and judging readiness for consequential work.
Any clinical, legal, engineering, financial, intelligence or safety-critical application.
Never do this from the page alone
Use practice forecasts to make live consequential decisions or present oneself as qualified.
Copy AI or supervisor advice before the independent forecast or hide overrides and conflicts.
2 · Get ready
Gather what you need and check the starting conditions
What you need
Qualified domain instructor, 20–40 representative resolved and de-identified cases and outcome key.
Probability scale, rationale sheet, calculator or scoring tool and calibration chart.
Declared domain, case-mix, conflict, AI/version and appeal record.
Before you start
Provider verifies scope, confidentiality, case representativeness and legitimate use.
Define the binary or categorical outcome and the resolution date before any forecast.
Agree that every case receives a probability, not only memorable successes.
3 · The method
Follow these steps in order
Lock domain and case mix
Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.
Why: Expertise can disappear outside its domain or on a selective case set.
Check: The case list matches the written inclusion rule.
Forecast independently
For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.
Why: Independent commitment preserves the learner’s judgement for calibration.
Check: A timestamped probability exists for every case.
Write diagnostic reasons
List the base rate, two case-specific cues and one reason the forecast could fail.
Why: Reasons expose whether confidence comes from relevant information or persuasive story.
Check: Each forecast has base rate, cues and counterargument.
Reveal and score
After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.
Why: A proper score rewards probability accuracy across all outcomes.
Check: Every valid case contributes to the declared score.
Compare with a simple benchmark
Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.
Why: Expert judgement adds value only if it can be separated from an appropriate baseline, not merely from guessing.
Check: The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.
Check calibration and discrimination
Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.
Why: Calibration and discrimination are distinct qualities.
Check: The chart shows both reliability by bin and case separation.
Repair and retest
The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.
Why: Specific feedback can improve a component without claiming generic expertise.
Check: The retest preserves case mix and documents one targeted repair.
4 · Worked example
See the whole method used once
Scenario
A qualified meteorological trainer gives a learner 20 archived, region-matched forecasts of measurable rain by the next day.
Walkthrough
They define rain as at least the official recorded threshold at the named station and exclude severe-weather operations.
Elena records a probability for every case before seeing outcome or model forecast.
For one case she uses a 30% base rate, moisture cues and the counterargument that the front may stall, then forecasts 60%.
Rain occurs on the 60% case, so she converts 60% to 0.60 and records (0.60−1)² = 0.16; she repeats that calculation for every resolved case.
Her 20-case mean is 0.21 versus 0.24 for the 30% base-rate benchmark; in her 70–89% bin the mean forecast is 0.78 but only 0.60 of cases occur.
The trainer identifies high-bin overconfidence, teaches a base-rate anchor and gives 20 fresh archived cases without calling the first difference proof of improvement.
Result
Elena can reproduce the Brier calculation, compare it on the same cases with the benchmark and identify a 0.18 high-bin calibration gap. Those descriptive results belong to this region, horizon and case mix and do not qualify her for live warnings or another domain.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
Right and wrong comparison
Moment
Right / safer
Wrong / riskier
Why it matters
Case selection
Use a declared representative resolved sample.
Choose famous cases where the expert was right.
Selection bias can create false expertise.
Commitment
Forecast before outcome, peer or AI view.
Adjust the remembered forecast after resolution.
Hindsight destroys calibration evidence.
Metric
Score all probabilities against outcomes.
Judge confidence, credentials or eloquent rationale.
Persuasion can diverge from accuracy.
Transfer
Retest the same domain and label new domains separately.
Assume calibrated weather judgement transfers to finance or medicine.
Expert performance is domain- and feedback-specific.
6 · Common mistakes
Spot the error and apply the correction
Common mistakes and corrections
Mistake
Fix
Only confident cases receive forecasts.
Require an entry for every sampled case.
Brier score is calculated after rounding probabilities to yes/no.
Use the original 0–1 probability for squared error.
Wide calibration bins hide overconfidence.
Predeclare bins and report counts in each.
AI advice contaminates the baseline.
Capture an independent forecast first and log any later model/version and update.
7 · Practice
Turn the steps into a usable skill
First session
Provider defines domain and resolution rule.
Complete five familiarisation cases.
Forecast a 20-case block independently.
Score calibration and discrimination.
Review one error class and set a matched retest.
Repeat plan
The provider schedules fresh resolved blocks until the sample is adequate and feedback remains representative. Live or operational cases are not used for unsupervised practice. Retest after delay and whenever the case mix, model or domain changes.
Progress when
Calibration improves across enough provider-approved cases.
Discrimination improves without cherry-picking or narrower case mix.
The learner states scope, uncertainty, conflicts and escalation accurately.
Do not progress when
Ground truth is ambiguous or outcomes are selectively missing.
Practice cases leak into live consequential decisions.
Improvement depends on seeing adviser or AI output before commitment.
8 · Check the result
Measure what changed
Provider-reviewed probabilistic accuracy on resolved domain cases.
How: For each resolved case calculate (p−o)² using forecast p on 0–1 and outcome o in {0,1}; average across cases for the Brier score. For each predeclared bin report mean p, occurred/total and their gap. Report mean p for occurred versus non-occurred cases, case count, base rate and missing outcomes.
Good result: A fresh matched block shows provider-defined improvement beyond sampling noise without narrowed case selection or hidden assistance.
This does not prove: It does not establish qualification, ethical authority, cross-domain expertise or safe live decision-making.
Self-check
What exact domain and case mix are scored?
Was every probability committed before feedback?
Can you distinguish calibration from discrimination?
Are conflicts and AI/version visible?
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
Stop if a training answer could affect a live person, asset or safety outcome.
Stop for confidentiality breach, unclear ground truth, conflict of interest or unlogged AI advice.
Route high-stakes judgements through the accountable qualified system, including override and appeal.
Accessibility and adaptations
Use plain-language cases, screen-reader-compatible probability controls and extended equal time.
Allow a scribe who records without advising.
Measure performance with declared supports rather than removing them.
10 · Evidence and limits
Why these instructions are here
primary research
Geopolitical forecasting research identified task-specific drivers of prediction accuracy and calibration rather than credentials alone.
Barbara Mellers; Eric Stone; Pavel Atanasov; Nick Rohrbaugh; S. Emlen Metz; Lyle Ungar; Michael M. Bishop; Michael Horowitz; Ed Merkle; Philip Tetlock · 2015 · Primary research
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
01
Lock domain and case mix
Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.
Why this step exists
Expertise can disappear outside its domain or on a selective case set.
Success check
The case list matches the written inclusion rule.
I’m stuck on this step
Reset: Re-read this authored instruction — “Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.” — and its success check, then attempt only this step.
Possible snag: Only confident cases receive forecasts.
Correction: Require an entry for every sampled case.
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
02
Forecast independently
For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.
Why this step exists
Independent commitment preserves the learner’s judgement for calibration.
Success check
A timestamped probability exists for every case.
I’m stuck on this step
Reset: Re-read this authored instruction — “For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.” — and its success check, then attempt only this step.
Possible snag: AI advice contaminates the baseline.
Correction: Capture an independent forecast first and log any later model/version and update.
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
03
Write diagnostic reasons
List the base rate, two case-specific cues and one reason the forecast could fail.
Why this step exists
Reasons expose whether confidence comes from relevant information or persuasive story.
Success check
Each forecast has base rate, cues and counterargument.
I’m stuck on this step
Reset: Re-read this authored instruction — “List the base rate, two case-specific cues and one reason the forecast could fail.” — and its success check, then attempt only this step.
Possible snag: The result from “List the base rate, two case-specific cues and one reason the forecast could fail.” does not yet meet this declared check: Each forecast has base rate, cues and counterargument.
Correction: Return to the start of “Write diagnostic reasons”, reduce complexity or pace, and repeat only the part needed to satisfy: “Each forecast has base rate, cues and counterargument.”
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
04
Reveal and score
After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.
Why this step exists
A proper score rewards probability accuracy across all outcomes.
Success check
Every valid case contributes to the declared score.
I’m stuck on this step
Reset: Re-read this authored instruction — “After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.” — and its success check, then attempt only this step.
Possible snag: Brier score is calculated after rounding probabilities to yes/no.
Correction: Use the original 0–1 probability for squared error.
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
05
Compare with a simple benchmark
Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.
Why this step exists
Expert judgement adds value only if it can be separated from an appropriate baseline, not merely from guessing.
Success check
The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.
I’m stuck on this step
Reset: Re-read this authored instruction — “Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.” — and its success check, then attempt only this step.
Possible snag: The result from “Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.” does not yet meet this declared check: The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.
Correction: Return to the start of “Compare with a simple benchmark”, reduce complexity or pace, and repeat only the part needed to satisfy: “The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.”
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
06
Check calibration and discrimination
Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.
Why this step exists
Calibration and discrimination are distinct qualities.
Success check
The chart shows both reliability by bin and case separation.
I’m stuck on this step
Reset: Re-read this authored instruction — “Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.” — and its success check, then attempt only this step.
Possible snag: Wide calibration bins hide overconfidence.
Correction: Predeclare bins and report counts in each.
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
07
Repair and retest
The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.
Why this step exists
Specific feedback can improve a component without claiming generic expertise.
Success check
The retest preserves case mix and documents one targeted repair.
I’m stuck on this step
Reset: Re-read this authored instruction — “The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.” — and its success check, then attempt only this step.
Possible snag: The result from “The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.” does not yet meet this declared check: The retest preserves case mix and documents one targeted repair.
Correction: Return to the start of “Repair and retest”, reduce complexity or pace, and repeat only the part needed to satisfy: “The retest preserves case mix and documents one targeted repair.”
Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Case selection — Selection bias can create false expertise.
Judge confidence, credentials or eloquent rationale.
Transfer — Expert performance is domain- and feedback-specific.
Correct / safer
Retest the same domain and label new domains separately.
Wrong / riskier
Assume calibrated weather judgement transfers to finance or medicine.
Method-structure checklist
10 of 10 structural checks present
✓ Ordered, Power-specific instructions — present
✓ Every activity has a success check — present
✓ Materials or supplied records are declared — present
✓ Measurement or assessment rule is present — present
✓ Tutorial-specific troubleshooting is present — present
✓ Stopping or escalation boundary is present — present
✓ Every activity has an adjacent alternative — present
✓ Correct-versus-incorrect comparison is present — present
✓ Evidence context is bound to the Power record — present
✓ Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.