Skip to content
TITAN//CAPABILITY
Revision 7T · Full Tutorial Edition · Updated 1 September 2026

PWR-118 · SUPERVISED full tutorial

Calibrate domain forecasts against resolved cases under expert supervision

Expert judgement is domain-specific performance in a feedback environment. A qualified domain expert selects representative resolved cases, protects sensitive information, defines the outcome and teaches probability forecasts. The learner commits a probability before feedback, then checks calibration—whether 70% forecasts occur about 70% of the time—and discrimination—whether higher probabilities went to events that happened. Credentials, confidence and eloquence are not substituted for resolved accuracy.

What you will produceUnder a qualified provider, the learner completes at least 20 resolved domain cases and produces a Brier score (mean squared probability error), calibration bands, discrimination and an error review for the declared case mix.
Method7 numbered Power-specific steps
Practice authorityFull method with qualified supervision where stated

One source of teaching truth

Full step-by-step individual tutorial · TLU-PWR-118

Under a qualified provider, the learner completes at least 20 resolved domain cases and produces a Brier score (mean squared probability error), calibration bands, discrimination and an error review for the declared case mix.

Canonical Power page
PWR-118 · Expert judgement
Full tutorial
Open full tutorial
Practical authority
The tutorial teaches the complete method; qualified supervision controls the specified practical parts.
Current treatment
Full supervised tutorial
Research depth
deep · 6 bound sources
Risk framing
high
Capability self-practice
Qualified supervision is required for the practical method
Pathway membership
G-CUR-011

Open My Power Path Inspect the canonical record

1 · Permission and limits

Know exactly what you may do

You may

  • State independent probability judgements, rationales, uncertainty and conflicts.
  • Use de-identified training cases selected by the provider.

Qualified help is required for

  • Defining domain scope and ground truth, selecting cases, protecting data, teaching standards and judging readiness for consequential work.
  • Any clinical, legal, engineering, financial, intelligence or safety-critical application.

Never do this from the page alone

  • Use practice forecasts to make live consequential decisions or present oneself as qualified.
  • Copy AI or supervisor advice before the independent forecast or hide overrides and conflicts.

2 · Get ready

Gather what you need and check the starting conditions

What you need

  • Qualified domain instructor, 20–40 representative resolved and de-identified cases and outcome key.
  • Probability scale, rationale sheet, calculator or scoring tool and calibration chart.
  • Declared domain, case-mix, conflict, AI/version and appeal record.

Before you start

  • Provider verifies scope, confidentiality, case representativeness and legitimate use.
  • Define the binary or categorical outcome and the resolution date before any forecast.
  • Agree that every case receives a probability, not only memorable successes.

3 · The method

Follow these steps in order

  1. Lock domain and case mix

    Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.

    Why: Expertise can disappear outside its domain or on a selective case set.

    Check: The case list matches the written inclusion rule.

  2. Forecast independently

    For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.

    Why: Independent commitment preserves the learner’s judgement for calibration.

    Check: A timestamped probability exists for every case.

  3. Write diagnostic reasons

    List the base rate, two case-specific cues and one reason the forecast could fail.

    Why: Reasons expose whether confidence comes from relevant information or persuasive story.

    Check: Each forecast has base rate, cues and counterargument.

  4. Reveal and score

    After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.

    Why: A proper score rewards probability accuracy across all outcomes.

    Check: Every valid case contributes to the declared score.

  5. Compare with a simple benchmark

    Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.

    Why: Expert judgement adds value only if it can be separated from an appropriate baseline, not merely from guessing.

    Check: The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.

  6. Check calibration and discrimination

    Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.

    Why: Calibration and discrimination are distinct qualities.

    Check: The chart shows both reliability by bin and case separation.

  7. Repair and retest

    The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.

    Why: Specific feedback can improve a component without claiming generic expertise.

    Check: The retest preserves case mix and documents one targeted repair.

4 · Worked example

See the whole method used once

Scenario

A qualified meteorological trainer gives a learner 20 archived, region-matched forecasts of measurable rain by the next day.

Walkthrough

  1. They define rain as at least the official recorded threshold at the named station and exclude severe-weather operations.
  2. Elena records a probability for every case before seeing outcome or model forecast.
  3. For one case she uses a 30% base rate, moisture cues and the counterargument that the front may stall, then forecasts 60%.
  4. Rain occurs on the 60% case, so she converts 60% to 0.60 and records (0.60−1)² = 0.16; she repeats that calculation for every resolved case.
  5. Her 20-case mean is 0.21 versus 0.24 for the 30% base-rate benchmark; in her 70–89% bin the mean forecast is 0.78 but only 0.60 of cases occur.
  6. The trainer identifies high-bin overconfidence, teaches a base-rate anchor and gives 20 fresh archived cases without calling the first difference proof of improvement.

Result

Elena can reproduce the Brier calculation, compare it on the same cases with the benchmark and identify a 0.18 high-bin calibration gap. Those descriptive results belong to this region, horizon and case mix and do not qualify her for live warnings or another domain.

5 · Right and wrong

Compare correct or safer execution with the common wrong version

Right and wrong comparison
MomentRight / saferWrong / riskierWhy it matters
Case selectionUse a declared representative resolved sample.Choose famous cases where the expert was right.Selection bias can create false expertise.
CommitmentForecast before outcome, peer or AI view.Adjust the remembered forecast after resolution.Hindsight destroys calibration evidence.
MetricScore all probabilities against outcomes.Judge confidence, credentials or eloquent rationale.Persuasion can diverge from accuracy.
TransferRetest the same domain and label new domains separately.Assume calibrated weather judgement transfers to finance or medicine.Expert performance is domain- and feedback-specific.

6 · Common mistakes

Spot the error and apply the correction

Common mistakes and corrections
MistakeFix
Only confident cases receive forecasts.Require an entry for every sampled case.
Brier score is calculated after rounding probabilities to yes/no.Use the original 0–1 probability for squared error.
Wide calibration bins hide overconfidence.Predeclare bins and report counts in each.
AI advice contaminates the baseline.Capture an independent forecast first and log any later model/version and update.

7 · Practice

Turn the steps into a usable skill

First session

  1. Provider defines domain and resolution rule.
  2. Complete five familiarisation cases.
  3. Forecast a 20-case block independently.
  4. Score calibration and discrimination.
  5. Review one error class and set a matched retest.

Repeat plan

The provider schedules fresh resolved blocks until the sample is adequate and feedback remains representative. Live or operational cases are not used for unsupervised practice. Retest after delay and whenever the case mix, model or domain changes.

Progress when

  • Calibration improves across enough provider-approved cases.
  • Discrimination improves without cherry-picking or narrower case mix.
  • The learner states scope, uncertainty, conflicts and escalation accurately.

Do not progress when

  • Ground truth is ambiguous or outcomes are selectively missing.
  • Practice cases leak into live consequential decisions.
  • Improvement depends on seeing adviser or AI output before commitment.

8 · Check the result

Measure what changed

Provider-reviewed probabilistic accuracy on resolved domain cases.

How: For each resolved case calculate (p−o)² using forecast p on 0–1 and outcome o in {0,1}; average across cases for the Brier score. For each predeclared bin report mean p, occurred/total and their gap. Report mean p for occurred versus non-occurred cases, case count, base rate and missing outcomes.

Good result: A fresh matched block shows provider-defined improvement beyond sampling noise without narrowed case selection or hidden assistance.

This does not prove: It does not establish qualification, ethical authority, cross-domain expertise or safe live decision-making.

Self-check

9 · Stop, adapt or get help

Keep the safety boundary practical

Stop and get help

Accessibility and adaptations

10 · Evidence and limits

Why these instructions are here

  1. primary research

    Geopolitical forecasting research identified task-specific drivers of prediction accuracy and calibration rather than credentials alone.

    The psychology of intelligence analysis: Drivers of prediction accuracy in world politics
  2. official guidance

    NIST AI RMF treats human–AI performance as a governed, measured configuration with risks, not as automatic expert enhancement.

    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

Limits

Open the complete canonical research register
  1. Primary empirical supportLimiting / contrary
    The psychology of intelligence analysis: Drivers of prediction accuracy in world politics

    Barbara Mellers; Eric Stone; Pavel Atanasov; Nick Rohrbaugh; S. Emlen Metz; Lyle Ungar; Michael M. Bishop; Michael Horowitz; Ed Merkle; Philip Tetlock · 2015 · Primary research

  2. Primary empirical supportLimiting / contrary
    To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making

    Zana Buçinca; Maja Barbara Malaya; Krzysztof Z. Gajos · 2021 · Primary research

  3. Limiting / contrary
    Do People Listen to Cassandra? Persuasion and Accuracy in Geopolitical Forecasts

    Yhonatan S. Shemesh; Avi Gamoran; David Leiser; Michael Gilead · 2025 · Primary research

  4. Primary empirical supportLimiting / contrary
    Debiasing Decisions

    Carey K. Morewedge; Haewon Yoon; Irene Scopelliti; Carl W. Symborski; James H. Korris; Karim S. Kassam · 2015 · Primary research

  5. Primary empirical supportLimiting / contrary
    The impact of AI errors in a human-in-the-loop process

    Ujué Agudo; Karlos G. Liberal; Miren Arrese; Helena Matute · 2024 · Primary research

  6. Limiting / contraryOfficial boundary context
    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    National Institute of Standards and Technology · 2023 · Official framework

Read the complete evidence interpretation on the Power dossier.

Tutorial delivery controls

Learn, adapt, troubleshoot and resume

Estimated timeEstimated 11 min reading; practical time is provider-set
DifficultyIntermediate
EquipmentBasic stationery or digital tools
SpaceDesk / seated
Method qualityComprehensive10 of 10 structural checks present. Automated method-readiness band; human editorial sign-off is separate.
Evidence contextG4; Deep research depthScientific support is evaluated separately from teaching-method structure.
Editorial reviewPending manual sign-offNo human approval is claimed until reviewer, date and content hash are recorded.
Your tutorial progress0 of 7 steps complete
0 of 7 steps complete
Download learner worksheet

Progress is saved only in this browser on this device.

Step-by-step learner mode

Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.

01

Lock domain and case mix

Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.

Why this step exists

Expertise can disappear outside its domain or on a selective case set.

Success check

The case list matches the written inclusion rule.

I’m stuck on this step

Reset: Re-read this authored instruction — “Name the exact task, population, time horizon and exclusions; the provider samples resolved cases from that frame.” — and its success check, then attempt only this step.

  1. Possible snag: Only confident cases receive forecasts.

    Correction: Require an entry for every sampled case.

Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.

02

Forecast independently

For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.

Why this step exists

Independent commitment preserves the learner’s judgement for calibration.

Success check

A timestamped probability exists for every case.

I’m stuck on this step

Reset: Re-read this authored instruction — “For each case, record a probability from 0 to 100 before seeing the outcome, peer view or AI advice.” — and its success check, then attempt only this step.

  1. Possible snag: AI advice contaminates the baseline.

    Correction: Capture an independent forecast first and log any later model/version and update.

Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.

03

Write diagnostic reasons

List the base rate, two case-specific cues and one reason the forecast could fail.

Why this step exists

Reasons expose whether confidence comes from relevant information or persuasive story.

Success check

Each forecast has base rate, cues and counterargument.

I’m stuck on this step

Reset: Re-read this authored instruction — “List the base rate, two case-specific cues and one reason the forecast could fail.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “List the base rate, two case-specific cues and one reason the forecast could fail.” does not yet meet this declared check: Each forecast has base rate, cues and counterargument.

    Correction: Return to the start of “Write diagnostic reasons”, reduce complexity or pace, and repeat only the part needed to satisfy: “Each forecast has base rate, cues and counterargument.”

Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.

04

Reveal and score

After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.

Why this step exists

A proper score rewards probability accuracy across all outcomes.

Success check

Every valid case contributes to the declared score.

I’m stuck on this step

Reset: Re-read this authored instruction — “After the block, reveal outcomes. Convert a forecast q% to p=q/100 and code the outcome o as 1 when it occurred or 0 when it did not. For each case calculate (p−o)²; the Brier score is the sum of those errors divided by the number of resolved cases.” — and its success check, then attempt only this step.

  1. Possible snag: Brier score is calculated after rounding probabilities to yes/no.

    Correction: Use the original 0–1 probability for squared error.

Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.

05

Compare with a simple benchmark

Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.

Why this step exists

Expert judgement adds value only if it can be separated from an appropriate baseline, not merely from guessing.

Success check

The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.

I’m stuck on this step

Reset: Re-read this authored instruction — “Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “Score the same cases using the historical base rate or another provider-approved simple rule, then compare it with the learner without changing the case set.” does not yet meet this declared check: The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.

    Correction: Return to the start of “Compare with a simple benchmark”, reduce complexity or pace, and repeat only the part needed to satisfy: “The record shows learner and benchmark scores on identical cases and names where the learner helped or hurt.”

Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.

06

Check calibration and discrimination

Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.

Why this step exists

Calibration and discrimination are distinct qualities.

Success check

The chart shows both reliability by bin and case separation.

I’m stuck on this step

Reset: Re-read this authored instruction — “Use the predeclared probability bands. In each band, average the forecast probabilities and separately divide occurred outcomes by resolved cases; their gap shows calibration. For discrimination, compare the mean forecast assigned to occurred cases with the mean assigned to non-occurred cases.” — and its success check, then attempt only this step.

  1. Possible snag: Wide calibration bins hide overconfidence.

    Correction: Predeclare bins and report counts in each.

Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.

07

Repair and retest

The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.

Why this step exists

Specific feedback can improve a component without claiming generic expertise.

Success check

The retest preserves case mix and documents one targeted repair.

I’m stuck on this step

Reset: Re-read this authored instruction — “The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “The provider classifies base-rate neglect, cue misuse, overconfidence or data leakage, teaches one correction and supplies a fresh matched case block.” does not yet meet this declared check: The retest preserves case mix and documents one targeted repair.

    Correction: Return to the start of “Repair and retest”, reduce complexity or pace, and repeat only the part needed to satisfy: “The retest preserves case mix and documents one targeted repair.”

Stop / get help: Stop if a training answer could affect a live person, asset or safety outcome.

Correct versus incorrect execution

These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.

Case selection — Selection bias can create false expertise.
PWR-118 correct and incorrect comparison: Case selectionCase selection. Correct or safer: Use a declared representative resolved sample.. Wrong or riskier: Choose famous cases where the expert was right.. Why: Selection bias can create false expertise.SITUATIONCase selectionCORRECT / SAFERUse a declared representative resolved sample.WRONG / RISKIERChoose famous cases where the expert was right.YESNO
Correct / safer

Use a declared representative resolved sample.

Wrong / riskier

Choose famous cases where the expert was right.

Commitment — Hindsight destroys calibration evidence.
PWR-118 correct and incorrect comparison: CommitmentCommitment. Correct or safer: Forecast before outcome, peer or AI view.. Wrong or riskier: Adjust the remembered forecast after resolution.. Why: Hindsight destroys calibration evidence.SITUATIONCommitmentCORRECT / SAFERForecast before outcome, peer or AI view.WRONG / RISKIERAdjust the remembered forecast after resolution.YESNO
Correct / safer

Forecast before outcome, peer or AI view.

Wrong / riskier

Adjust the remembered forecast after resolution.

Metric — Persuasion can diverge from accuracy.
PWR-118 correct and incorrect comparison: MetricMetric. Correct or safer: Score all probabilities against outcomes.. Wrong or riskier: Judge confidence, credentials or eloquent rationale.. Why: Persuasion can diverge from accuracy.SITUATIONMetricCORRECT / SAFERScore all probabilities against outcomes.WRONG / RISKIERJudge confidence, credentials or eloquentrationale.YESNO
Correct / safer

Score all probabilities against outcomes.

Wrong / riskier

Judge confidence, credentials or eloquent rationale.

Transfer — Expert performance is domain- and feedback-specific.
PWR-118 correct and incorrect comparison: TransferTransfer. Correct or safer: Retest the same domain and label new domains separately.. Wrong or riskier: Assume calibrated weather judgement transfers to finance or medicine.. Why: Expert performance is domain- and feedback-specific.SITUATIONTransferCORRECT / SAFERRetest the same domain and label new domainsseparately.WRONG / RISKIERAssume calibrated weather judgement transfers tofinance or medicine.YESNO
Correct / safer

Retest the same domain and label new domains separately.

Wrong / riskier

Assume calibrated weather judgement transfers to finance or medicine.

Method-structure checklist

10 of 10 structural checks present

  • Ordered, Power-specific instructions — present
  • Every activity has a success check — present
  • Materials or supplied records are declared — present
  • Measurement or assessment rule is present — present
  • Tutorial-specific troubleshooting is present — present
  • Stopping or escalation boundary is present — present
  • Every activity has an adjacent alternative — present
  • Correct-versus-incorrect comparison is present — present
  • Evidence context is bound to the Power record — present
  • Planning metadata is present — present

The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.

Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.