Skip to content
TITAN//CAPABILITY
Revision 7T · Full Tutorial Edition · Updated 1 September 2026

PWR-224 · RESEARCH full tutorial

Evaluate a predictive capability model without turning a narrow forecast into a destiny score

This research tutorial teaches a complete predictive-model audit. You will define one outcome with a forecast horizon; inspect target population and data splits; compare calibration with discrimination; check subgroups, uncertainty, drift, abstention and decision consequences; then write a bounded evidence grade. No personal ranking or consequential prediction is performed.

What you will produceThe learner produces a reproducible model-card audit that states what is predicted, for whom and when; reports validation plus subgroup limits; and rejects unsupported claims about worth, potential or destiny.
Method10 numbered Power-specific steps
Practice authorityComplete research method; no operational self-experiment

One source of teaching truth

Full step-by-step individual tutorial · TLU-PWR-224

The learner produces a reproducible model-card audit that states what is predicted, for whom and when; reports validation plus subgroup limits; and rejects unsupported claims about worth, potential or destiny.

Canonical Power page
PWR-224 · Predictive capability modelling
Full tutorial
Open full tutorial
Practical authority
The tutorial teaches a complete research method; operational self-experiment is excluded.
Current treatment
Full research tutorial
Research depth
deep · 7 bound sources
Risk framing
critical
Capability self-practice
Research method only; no operational self-experiment
Pathway membership
G-CUR-022

Open My Power Path Inspect the canonical record

1 · Permission and limits

Know exactly what you may do

You may

  • Audit published aggregate model results or a fictional dataset.
  • Calculate simple calibration by comparing predicted and observed frequencies in supplied bins.
  • Recommend abstention, external validation or retirement when evidence is insufficient.

Qualified help is required for

  • Building or deploying a real model that affects education, employment, insurance, benefits, health, credit or liberty.
  • Access to personal data, fairness assessment for protected groups or impact evaluation.
  • Approval of thresholds, human override, appeals, monitoring and model changes in a real institution.

Never do this from the page alone

  • Score a real person’s potential, worth or future from this lesson.
  • Use a model trained for one outcome, population or horizon to decide another.
  • Hide subgroup error, uncertainty or abstention behind one overall accuracy number.

2 · Get ready

Gather what you need and check the starting conditions

What you need

  • A supplied fictional model report predicting whether 100 fictional trainees finish a six-week optional course. In the 2025 test set, 75 of 100 completed, the model classified 78 correctly and an always-complete base-rate rule classified 75 correctly.
  • Four supplied calibration rows: predicted 50%, n=20, 10 completed; predicted 70%, n=20, 14 completed; predicted 80%, n=20, 12 completed; predicted 90%, n=40, 39 completed. The accessibility subgroup has n=8 and its outcomes are suppressed. A 2026 drift snapshot has 60 completions among 100 people and no recalibration.
  • Audit worksheet, calculator and decision-consequence table.

Before you start

  • Confirm that every row is fictional or aggregate and no real person can be identified.
  • Write the outcome, horizon, target population, allowed research question and prohibited uses.
  • Predeclare the minimum evidence: separated test set, comparator, calibration, subgroup reporting, uncertainty and monitoring.

3 · The method

Follow these steps in order

  1. Lock outcome and horizon

    Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.

    Why: Outcome blur allows a score to be reused for unrelated decisions.

    Check: The target can be marked observed or not observed at one declared time.

  2. Define intended population

    Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.

    Why: Performance depends on who generated examples.

    Check: The audit names target population and every group to which this result must not transfer.

  3. Inspect data separation

    Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.

    Why: Leakage can create impressive but unusable performance.

    Check: The final test is plausibly independent, or available evidence grade is stopped.

  4. Compare a simple baseline

    Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.

    Why: Extra complexity is useful only when it improves the declared forecast against a fair comparator.

    Check: Any claimed gain is shown for identical cases and outcome.

  5. Separate discrimination from calibration

    Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.

    Why: A model can rank correctly while giving misleading probabilities.

    Check: Both ranking quality and probability accuracy are stated in plain English.

  6. Inspect subgroup and missing-data error

    Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.

    Why: An average can distribute errors unfairly or conceal exclusion.

    Check: The audit marks every subgroup as tested, underpowered or absent.

  7. Check uncertainty and abstention

    Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.

    Why: Forced predictions turn lack of knowledge into false precision.

    Check: Every uncertain case has a non-punitive route that does not default to denial.

  8. Test drift and versioning

    Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.

    Why: Performance can decay when people, policy or measurement changes.

    Check: The audit names this model version, monitoring interval and stop threshold.

  9. Map decision consequences

    For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.

    Why: Prediction quality is not identical as decision value or legitimacy.

    Check: Worksheet separates forecast error from action harm.

  10. Write evidence grade

    Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.

    Why: A bounded grade prevents a narrow result becoming a destiny claim.

    Check: The conclusion includes context, uncertainty, groups, consequences and prohibited uses.

4 · Worked example

See the whole method used once

Scenario

A fictional model claims to predict six-week course completion with 78% accuracy and is advertised as showing each trainee’s “future potential.”

Walkthrough

  1. The learner rewrites the target outcome as completion of this optional course within six weeks in a fictional 2025 cohort.
  2. They discover that 75% completed anyway, so this model’s three-point accuracy gain over the base-rate comparator is small.
  3. The calibration table shows that only 60% of people in the 80% prediction bin completed; a small accessibility subgroup is not reported.
  4. A 2026 snapshot has a different base rate. The learner grades this model for internal research only and prohibits admission, employment or personal-potential use.

Result

The audit preserves a narrow course-completion signal but rejects calibrated-probability, subgroup, drift, benefit and destiny claims.

5 · Right and wrong

Compare correct or safer execution with the common wrong version

Right and wrong comparison
MomentRight / saferWrong / riskierWhy it matters
Reading accuracyCompare it with identical-population base rate and error costs.Call 78% impressive without a comparator.A high number may barely beat a trivial rule.
Reading probabilityCheck observed frequencies in probability bins.Treat a rank score as calibrated individual probability.Discrimination and calibration answer different questions; one cannot stand in for the other.
Handling missing groupsMark transportability and fairness unresolved when a relevant group is absent.Assume overall result applies equally to everyone.Unreported groups cannot inherit average validity.
Using uncertaintyAllow abstention and provide a human review plus correction route.Force every case into a rank.Forced certainty magnifies error when a case falls outside available evidence.

6 · Common mistakes

Spot the error and apply the correction

Common mistakes and corrections
MistakeFix
The course-completion forecast is relabelled as a measure of future potential.Rewrite this result with one outcome, horizon, population and model version every time it is cited.
The report quotes accuracy without a baseline comparator.Add a transparent base-rate or existing-process comparison on identical test cases.
The audit treats a ranking score as an individual probability.Build predicted-versus-observed bins or mark probability claims unsupported.
Only favourable subgroup results are reported after looking at source data.Predeclare relevant groups and report all of them, including small or missing samples, without cherry-picking.
A later cohort is scored without checking drift or this model version.Compare base rates and calibration over time, then set a withdrawal threshold.

7 · Practice

Turn the steps into a usable skill

First session

  1. Predeclare the fictional outcome, population, evidence gate and prohibited uses.
  2. Check the splits, leakage and base-rate comparator.
  3. Calculate the supplied calibration bins plus error types.
  4. Audit the subgroup, missing-data, uncertainty and drift fields.
  5. Map the decision consequences, then write a bounded evidence grade.

Repeat plan

Audit one new fictional or published aggregate model each week for four weeks. Alternate a well-calibrated example with a poorly calibrated one. Progress only to decision-impact studies, never to personal scoring.

Progress when

  • Always state target outcome, horizon, population, version and comparator before quoting a score.
  • Calibration, subgroup errors, uncertainty and drift are interpreted correctly.
  • The final grade stays within the available validation and decision-impact evidence.

Do not progress when

  • Data separation, target definition or model version cannot be established.
  • Exercise shifts to a real person or a consequential decision.
  • A missing subgroup or drift failure is being treated as proof of no problem.

8 · Check the result

Measure what changed

Completeness and correctness of a predictive-model evidence audit

How: Score ten fields—outcome/horizon, population, split, comparator, discrimination, calibration, subgroups, uncertainty, drift and consequences—with traceable values plus one evidence grade.

Good result: All fields are correct or explicitly unresolved, this grade matches the weakest essential field and no personal-potential claim remains.

This does not prove: It does not validate a real model, establish causal benefit, prove fairness or authorise any consequential decision.

Self-check

9 · Stop, adapt or get help

Keep the safety boundary practical

Stop and get help

Accessibility and adaptations

10 · Evidence and limits

Why these instructions are here

  1. official guidance

    NIST’s AI Risk Management Framework calls for context, validity, reliability, transparency and risk governance rather than treating a model score as universal truth.

    Artificial Intelligence Risk Management Framework (AI RMF 1.0)
  2. primary research

    A prospective digital-twin prediction study reported variable agreement and implementation errors, illustrating the limits of transporting model outputs across contexts.

    Development and Verification of a Digital Twin Patient Model to Predict Specific Treatment Response During the First 24 Hours of Sepsis

Limits

Open the complete canonical research register
  1. Primary empirical supportLimiting / contrary
    Development and Verification of a Digital Twin Patient Model to Predict Specific Treatment Response During the First 24 Hours of Sepsis

    A Lal; G Li; E Cubro; S Chalmers; H Li; V Herasevich; Y Dong; B W Pickering; O Kilickaya; O Gajic · 2020 · Primary research

  2. Limiting / contraryOfficial boundary context
    Artificial Intelligence Risk Management Framework (AI RMF 1.0)

    Elham Tabassi; National Institute of Standards and Technology · 2023 · Official standard

  3. Limiting / contraryOfficial boundary context
    NIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0

    National Institute of Standards and Technology · 2020 · Official standard

  4. Limiting / contraryOfficial boundary context
    Credibility of Computational Models Program: Research on Computational Models and Simulation Associated with Medical Devices

    U.S. Food and Drug Administration, Office of Science and Engineering Laboratories · 2023 · Official guidance

  5. Limiting / contraryOfficial boundary context
    Clinical Decision Support Software

    U.S. Food and Drug Administration · 2026 · Official guidance

  6. Limiting / contraryOfficial boundary context
    Cybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions

    United States Food and Drug Administration · 2026 · Official guidance

  7. Limiting / contraryOfficial boundary context
    Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions

    U.S. Food and Drug Administration · 2025 · Official guidance

Read the complete evidence interpretation on the Power dossier.

Tutorial delivery controls

Learn, adapt, troubleshoot and resume

Estimated timeEstimated 31 min reading and worksheet pass
DifficultyAdvanced
EquipmentCommon household or practice equipment
SpaceDesk / seated
Method qualityComprehensive10 of 10 structural checks present. Automated method-readiness band; human editorial sign-off is separate.
Evidence contextG5; Deep research depthScientific support is evaluated separately from teaching-method structure.
Editorial reviewPending manual sign-offNo human approval is claimed until reviewer, date and content hash are recorded.
Your tutorial progress0 of 10 steps complete
0 of 10 steps complete
Download learner worksheet

Progress is saved only in this browser on this device.

Step-by-step learner mode

Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.

01

Lock outcome and horizon

Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.

Why this step exists

Outcome blur allows a score to be reused for unrelated decisions.

Success check

The target can be marked observed or not observed at one declared time.

I’m stuck on this step

Reset: Re-read this authored instruction — “Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.” — and its success check, then attempt only this step.

  1. Possible snag: The course-completion forecast is relabelled as a measure of future potential.

    Correction: Rewrite this result with one outcome, horizon, population and model version every time it is cited.

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

02

Define intended population

Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.

Why this step exists

Performance depends on who generated examples.

Success check

The audit names target population and every group to which this result must not transfer.

I’m stuck on this step

Reset: Re-read this authored instruction — “Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.” — and its success check, then attempt only this step.

  1. Possible snag: Only favourable subgroup results are reported after looking at source data.

    Correction: Predeclare relevant groups and report all of them, including small or missing samples, without cherry-picking.

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

03

Inspect data separation

Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.

Why this step exists

Leakage can create impressive but unusable performance.

Success check

The final test is plausibly independent, or available evidence grade is stopped.

I’m stuck on this step

Reset: Re-read this authored instruction — “Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.” does not yet meet this declared check: The final test is plausibly independent, or available evidence grade is stopped.

    Correction: Return to the start of “Inspect data separation”, reduce complexity or pace, and repeat only the part needed to satisfy: “The final test is plausibly independent, or available evidence grade is stopped.”

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

04

Compare a simple baseline

Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.

Why this step exists

Extra complexity is useful only when it improves the declared forecast against a fair comparator.

Success check

Any claimed gain is shown for identical cases and outcome.

I’m stuck on this step

Reset: Re-read this authored instruction — “Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.” — and its success check, then attempt only this step.

  1. Possible snag: The report quotes accuracy without a baseline comparator.

    Correction: Add a transparent base-rate or existing-process comparison on identical test cases.

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

05

Separate discrimination from calibration

Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.

Why this step exists

A model can rank correctly while giving misleading probabilities.

Success check

Both ranking quality and probability accuracy are stated in plain English.

I’m stuck on this step

Reset: Re-read this authored instruction — “Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.” — and its success check, then attempt only this step.

  1. Possible snag: The audit treats a ranking score as an individual probability.

    Correction: Build predicted-versus-observed bins or mark probability claims unsupported.

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

06

Inspect subgroup and missing-data error

Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.

Why this step exists

An average can distribute errors unfairly or conceal exclusion.

Success check

The audit marks every subgroup as tested, underpowered or absent.

I’m stuck on this step

Reset: Re-read this authored instruction — “Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.” does not yet meet this declared check: The audit marks every subgroup as tested, underpowered or absent.

    Correction: Return to the start of “Inspect subgroup and missing-data error”, reduce complexity or pace, and repeat only the part needed to satisfy: “The audit marks every subgroup as tested, underpowered or absent.”

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

07

Check uncertainty and abstention

Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.

Why this step exists

Forced predictions turn lack of knowledge into false precision.

Success check

Every uncertain case has a non-punitive route that does not default to denial.

I’m stuck on this step

Reset: Re-read this authored instruction — “Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.” does not yet meet this declared check: Every uncertain case has a non-punitive route that does not default to denial.

    Correction: Return to the start of “Check uncertainty and abstention”, reduce complexity or pace, and repeat only the part needed to satisfy: “Every uncertain case has a non-punitive route that does not default to denial.”

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

08

Test drift and versioning

Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.

Why this step exists

Performance can decay when people, policy or measurement changes.

Success check

The audit names this model version, monitoring interval and stop threshold.

I’m stuck on this step

Reset: Re-read this authored instruction — “Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.” — and its success check, then attempt only this step.

  1. Possible snag: A later cohort is scored without checking drift or this model version.

    Correction: Compare base rates and calibration over time, then set a withdrawal threshold.

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

09

Map decision consequences

For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.

Why this step exists

Prediction quality is not identical as decision value or legitimacy.

Success check

Worksheet separates forecast error from action harm.

I’m stuck on this step

Reset: Re-read this authored instruction — “For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.” does not yet meet this declared check: Worksheet separates forecast error from action harm.

    Correction: Return to the start of “Map decision consequences”, reduce complexity or pace, and repeat only the part needed to satisfy: “Worksheet separates forecast error from action harm.”

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

10

Write evidence grade

Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.

Why this step exists

A bounded grade prevents a narrow result becoming a destiny claim.

Success check

The conclusion includes context, uncertainty, groups, consequences and prohibited uses.

I’m stuck on this step

Reset: Re-read this authored instruction — “Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.” — and its success check, then attempt only this step.

  1. Possible snag: The result from “Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.” does not yet meet this declared check: The conclusion includes context, uncertainty, groups, consequences and prohibited uses.

    Correction: Return to the start of “Write evidence grade”, reduce complexity or pace, and repeat only the part needed to satisfy: “The conclusion includes context, uncertainty, groups, consequences and prohibited uses.”

Stop / get help: Stop if real personal or protected data is introduced without authorised governance.

Correct versus incorrect execution

These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.

Reading accuracy — A high number may barely beat a trivial rule.
PWR-224 correct and incorrect comparison: Reading accuracyReading accuracy. Correct or safer: Compare it with identical-population base rate and error costs.. Wrong or riskier: Call 78% impressive without a comparator.. Why: A high number may barely beat a trivial rule.SITUATIONReading accuracyCORRECT / SAFERCompare it with identical-population base rateand error costs.WRONG / RISKIERCall 78% impressive without a comparator.YESNO
Correct / safer

Compare it with identical-population base rate and error costs.

Wrong / riskier

Call 78% impressive without a comparator.

Reading probability — Discrimination and calibration answer different questions; one cannot stand in for the other.
PWR-224 correct and incorrect comparison: Reading probabilityReading probability. Correct or safer: Check observed frequencies in probability bins.. Wrong or riskier: Treat a rank score as calibrated individual probability.. Why: Discrimination and calibration answer different questions; one cannot stand in for the other.SITUATIONReading probabilityCORRECT / SAFERCheck observed frequencies in probability bins.WRONG / RISKIERTreat a rank score as calibrated individualprobability.YESNO
Correct / safer

Check observed frequencies in probability bins.

Wrong / riskier

Treat a rank score as calibrated individual probability.

Handling missing groups — Unreported groups cannot inherit average validity.
PWR-224 correct and incorrect comparison: Handling missing groupsHandling missing groups. Correct or safer: Mark transportability and fairness unresolved when a relevant group is absent.. Wrong or riskier: Assume overall result applies equally to everyone.. Why: Unreported groups cannot inherit average validity.SITUATIONHandling missinggroupsCORRECT / SAFERMark transportability and fairness unresolvedwhen a relevant group is absent.WRONG / RISKIERAssume overall result applies equally toeveryone.YESNO
Correct / safer

Mark transportability and fairness unresolved when a relevant group is absent.

Wrong / riskier

Assume overall result applies equally to everyone.

Using uncertainty — Forced certainty magnifies error when a case falls outside available evidence.
PWR-224 correct and incorrect comparison: Using uncertaintyUsing uncertainty. Correct or safer: Allow abstention and provide a human review plus correction route.. Wrong or riskier: Force every case into a rank.. Why: Forced certainty magnifies error when a case falls outside available evidence.SITUATIONUsing uncertaintyCORRECT / SAFERAllow abstention and provide a human review pluscorrection route.WRONG / RISKIERForce every case into a rank.YESNO
Correct / safer

Allow abstention and provide a human review plus correction route.

Wrong / riskier

Force every case into a rank.

Method-structure checklist

10 of 10 structural checks present

  • Ordered, Power-specific instructions — present
  • Every activity has a success check — present
  • Materials or supplied records are declared — present
  • Measurement or assessment rule is present — present
  • Tutorial-specific troubleshooting is present — present
  • Stopping or escalation boundary is present — present
  • Every activity has an adjacent alternative — present
  • Correct-versus-incorrect comparison is present — present
  • Evidence context is bound to the Power record — present
  • Planning metadata is present — present

The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.

Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.