Section 406 of 440
Complete canonical tutorial. This reader section contains the same teaching body as PWR-224 · Predictive capability modelling. Open the Power dossier.
PWR-224 · RESEARCH full tutorial
Evaluate a predictive capability model without turning a narrow forecast into a destiny score
This research tutorial teaches a complete predictive-model audit. You will define one outcome with a forecast horizon; inspect target population and data splits; compare calibration with discrimination; check subgroups, uncertainty, drift, abstention and decision consequences; then write a bounded evidence grade. No personal ranking or consequential prediction is performed.
1 · Permission and limits
Know exactly what you may do
2 · Get ready
Gather what you need and check the starting conditions
What you need
- A supplied fictional model report predicting whether 100 fictional trainees finish a six-week optional course. In the 2025 test set, 75 of 100 completed, the model classified 78 correctly and an always-complete base-rate rule classified 75 correctly.
- Four supplied calibration rows: predicted 50%, n=20, 10 completed; predicted 70%, n=20, 14 completed; predicted 80%, n=20, 12 completed; predicted 90%, n=40, 39 completed. The accessibility subgroup has n=8 and its outcomes are suppressed. A 2026 drift snapshot has 60 completions among 100 people and no recalibration.
- Audit worksheet, calculator and decision-consequence table.
Before you start
- Confirm that every row is fictional or aggregate and no real person can be identified.
- Write the outcome, horizon, target population, allowed research question and prohibited uses.
- Predeclare the minimum evidence: separated test set, comparator, calibration, subgroup reporting, uncertainty and monitoring.
3 · The method
Follow these steps in order
- Lock outcome and horizon
Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.
Why: Outcome blur allows a score to be reused for unrelated decisions.
Check: The target can be marked observed or not observed at one declared time.
- Define intended population
Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.
Why: Performance depends on who generated examples.
Check: The audit names target population and every group to which this result must not transfer.
- Inspect data separation
Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.
Why: Leakage can create impressive but unusable performance.
Check: The final test is plausibly independent, or available evidence grade is stopped.
- Compare a simple baseline
Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.
Why: Extra complexity is useful only when it improves the declared forecast against a fair comparator.
Check: Any claimed gain is shown for identical cases and outcome.
- Separate discrimination from calibration
Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.
Why: A model can rank correctly while giving misleading probabilities.
Check: Both ranking quality and probability accuracy are stated in plain English.
- Inspect subgroup and missing-data error
Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.
Why: An average can distribute errors unfairly or conceal exclusion.
Check: The audit marks every subgroup as tested, underpowered or absent.
- Check uncertainty and abstention
Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.
Why: Forced predictions turn lack of knowledge into false precision.
Check: Every uncertain case has a non-punitive route that does not default to denial.
- Test drift and versioning
Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.
Why: Performance can decay when people, policy or measurement changes.
Check: The audit names this model version, monitoring interval and stop threshold.
- Map decision consequences
For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.
Why: Prediction quality is not identical as decision value or legitimacy.
Check: Worksheet separates forecast error from action harm.
- Write evidence grade
Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.
Why: A bounded grade prevents a narrow result becoming a destiny claim.
Check: The conclusion includes context, uncertainty, groups, consequences and prohibited uses.
4 · Worked example
See the whole method used once
Scenario
A fictional model claims to predict six-week course completion with 78% accuracy and is advertised as showing each trainee’s “future potential.”
Walkthrough
- The learner rewrites the target outcome as completion of this optional course within six weeks in a fictional 2025 cohort.
- They discover that 75% completed anyway, so this model’s three-point accuracy gain over the base-rate comparator is small.
- The calibration table shows that only 60% of people in the 80% prediction bin completed; a small accessibility subgroup is not reported.
- A 2026 snapshot has a different base rate. The learner grades this model for internal research only and prohibits admission, employment or personal-potential use.
Result
The audit preserves a narrow course-completion signal but rejects calibrated-probability, subgroup, drift, benefit and destiny claims.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
| Moment | Right / safer | Wrong / riskier | Why it matters |
|---|---|---|---|
| Reading accuracy | Compare it with identical-population base rate and error costs. | Call 78% impressive without a comparator. | A high number may barely beat a trivial rule. |
| Reading probability | Check observed frequencies in probability bins. | Treat a rank score as calibrated individual probability. | Discrimination and calibration answer different questions; one cannot stand in for the other. |
| Handling missing groups | Mark transportability and fairness unresolved when a relevant group is absent. | Assume overall result applies equally to everyone. | Unreported groups cannot inherit average validity. |
| Using uncertainty | Allow abstention and provide a human review plus correction route. | Force every case into a rank. | Forced certainty magnifies error when a case falls outside available evidence. |
6 · Common mistakes
Spot the error and apply the correction
| Mistake | Fix |
|---|---|
| The course-completion forecast is relabelled as a measure of future potential. | Rewrite this result with one outcome, horizon, population and model version every time it is cited. |
| The report quotes accuracy without a baseline comparator. | Add a transparent base-rate or existing-process comparison on identical test cases. |
| The audit treats a ranking score as an individual probability. | Build predicted-versus-observed bins or mark probability claims unsupported. |
| Only favourable subgroup results are reported after looking at source data. | Predeclare relevant groups and report all of them, including small or missing samples, without cherry-picking. |
| A later cohort is scored without checking drift or this model version. | Compare base rates and calibration over time, then set a withdrawal threshold. |
7 · Practice
Turn the steps into a usable skill
First session
- Predeclare the fictional outcome, population, evidence gate and prohibited uses.
- Check the splits, leakage and base-rate comparator.
- Calculate the supplied calibration bins plus error types.
- Audit the subgroup, missing-data, uncertainty and drift fields.
- Map the decision consequences, then write a bounded evidence grade.
Repeat plan
Audit one new fictional or published aggregate model each week for four weeks. Alternate a well-calibrated example with a poorly calibrated one. Progress only to decision-impact studies, never to personal scoring.
Progress when
- Always state target outcome, horizon, population, version and comparator before quoting a score.
- Calibration, subgroup errors, uncertainty and drift are interpreted correctly.
- The final grade stays within the available validation and decision-impact evidence.
Do not progress when
- Data separation, target definition or model version cannot be established.
- Exercise shifts to a real person or a consequential decision.
- A missing subgroup or drift failure is being treated as proof of no problem.
8 · Check the result
Measure what changed
Completeness and correctness of a predictive-model evidence audit
How: Score ten fields—outcome/horizon, population, split, comparator, discrimination, calibration, subgroups, uncertainty, drift and consequences—with traceable values plus one evidence grade.
Good result: All fields are correct or explicitly unresolved, this grade matches the weakest essential field and no personal-potential claim remains.
This does not prove: It does not validate a real model, establish causal benefit, prove fairness or authorise any consequential decision.
Self-check
- What trivial comparator must this model beat?
- Can you explain the difference between discrimination and calibration?
- Which group or context is absent, what claim must therefore stop?
- What happens to a person when this model abstains or makes each error type?
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
- Stop if real personal or protected data is introduced without authorised governance.
- Stop if output is proposed for admission, employment, insurance, credit, health, benefits, discipline or another consequential decision.
- Route model deployment, legal, privacy, fairness, accessibility and appeal design to accountable specialists and affected people.
Accessibility and adaptations
- Explain accuracy, calibration and false-positive trade-offs with frequencies and plain-language tables rather than formulas alone.
- Provide screen-reader-friendly tables plus text descriptions of reliability plots.
- Allow calculators and worked examples; do not equate slower arithmetic with weaker model judgement.
10 · Evidence and limits
Why these instructions are here
- official guidance
NIST’s AI Risk Management Framework calls for context, validity, reliability, transparency and risk governance rather than treating a model score as universal truth.
Artificial Intelligence Risk Management Framework (AI RMF 1.0) - primary research
A prospective digital-twin prediction study reported variable agreement and implementation errors, illustrating the limits of transporting model outputs across contexts.
Development and Verification of a Digital Twin Patient Model to Predict Specific Treatment Response During the First 24 Hours of Sepsis
Limits
- This tutorial evaluates evidence; it does not build, tune or deploy a model.
- A prediction is tied to one outcome, horizon, population, dataset and version.
- No model score measures human worth, destiny or total capability.
- A context-specific forecast cannot measure worth, destiny or total capability.
- The prospective digital-twin study showed variable agreement and common implementation errors; no cross-context universal capability prediction or improved decision outcome was established.
Open the complete canonical research register
- Primary empirical supportLimiting / contraryDevelopment and Verification of a Digital Twin Patient Model to Predict Specific Treatment Response During the First 24 Hours of Sepsis
A Lal; G Li; E Cubro; S Chalmers; H Li; V Herasevich; Y Dong; B W Pickering; O Kilickaya; O Gajic · 2020 · Primary research
- Limiting / contraryOfficial boundary contextArtificial Intelligence Risk Management Framework (AI RMF 1.0)
Elham Tabassi; National Institute of Standards and Technology · 2023 · Official standard
- Limiting / contraryOfficial boundary contextNIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0
National Institute of Standards and Technology · 2020 · Official standard
- Limiting / contraryOfficial boundary contextCredibility of Computational Models Program: Research on Computational Models and Simulation Associated with Medical Devices
U.S. Food and Drug Administration, Office of Science and Engineering Laboratories · 2023 · Official guidance
- Limiting / contraryOfficial boundary contextClinical Decision Support Software
U.S. Food and Drug Administration · 2026 · Official guidance
- Limiting / contraryOfficial boundary contextCybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions
United States Food and Drug Administration · 2026 · Official guidance
- Limiting / contraryOfficial boundary contextMarketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions
U.S. Food and Drug Administration · 2025 · Official guidance
Read the complete evidence interpretation on the Power dossier.
Tutorial delivery controls
Learn, adapt, troubleshoot and resume
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
Lock outcome and horizon
Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.
Outcome blur allows a score to be reused for unrelated decisions.
The target can be marked observed or not observed at one declared time.
I’m stuck on this step
Reset: Re-read this authored instruction — “Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.” — and its success check, then attempt only this step.
Possible snag: The course-completion forecast is relabelled as a measure of future potential.
Correction: Rewrite this result with one outcome, horizon, population and model version every time it is cited.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Define intended population
Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.
Performance depends on who generated examples.
The audit names target population and every group to which this result must not transfer.
I’m stuck on this step
Reset: Re-read this authored instruction — “Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.” — and its success check, then attempt only this step.
Possible snag: Only favourable subgroup results are reported after looking at source data.
Correction: Predeclare relevant groups and report all of them, including small or missing samples, without cherry-picking.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Inspect data separation
Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.
Leakage can create impressive but unusable performance.
The final test is plausibly independent, or available evidence grade is stopped.
I’m stuck on this step
Reset: Re-read this authored instruction — “Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.” — and its success check, then attempt only this step.
Possible snag: The result from “Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.” does not yet meet this declared check: The final test is plausibly independent, or available evidence grade is stopped.
Correction: Return to the start of “Inspect data separation”, reduce complexity or pace, and repeat only the part needed to satisfy: “The final test is plausibly independent, or available evidence grade is stopped.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Compare a simple baseline
Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.
Extra complexity is useful only when it improves the declared forecast against a fair comparator.
Any claimed gain is shown for identical cases and outcome.
I’m stuck on this step
Reset: Re-read this authored instruction — “Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.” — and its success check, then attempt only this step.
Possible snag: The report quotes accuracy without a baseline comparator.
Correction: Add a transparent base-rate or existing-process comparison on identical test cases.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Separate discrimination from calibration
Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.
A model can rank correctly while giving misleading probabilities.
Both ranking quality and probability accuracy are stated in plain English.
I’m stuck on this step
Reset: Re-read this authored instruction — “Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.” — and its success check, then attempt only this step.
Possible snag: The audit treats a ranking score as an individual probability.
Correction: Build predicted-versus-observed bins or mark probability claims unsupported.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Inspect subgroup and missing-data error
Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.
An average can distribute errors unfairly or conceal exclusion.
The audit marks every subgroup as tested, underpowered or absent.
I’m stuck on this step
Reset: Re-read this authored instruction — “Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.” — and its success check, then attempt only this step.
Possible snag: The result from “Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.” does not yet meet this declared check: The audit marks every subgroup as tested, underpowered or absent.
Correction: Return to the start of “Inspect subgroup and missing-data error”, reduce complexity or pace, and repeat only the part needed to satisfy: “The audit marks every subgroup as tested, underpowered or absent.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Check uncertainty and abstention
Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.
Forced predictions turn lack of knowledge into false precision.
Every uncertain case has a non-punitive route that does not default to denial.
I’m stuck on this step
Reset: Re-read this authored instruction — “Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.” — and its success check, then attempt only this step.
Possible snag: The result from “Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.” does not yet meet this declared check: Every uncertain case has a non-punitive route that does not default to denial.
Correction: Return to the start of “Check uncertainty and abstention”, reduce complexity or pace, and repeat only the part needed to satisfy: “Every uncertain case has a non-punitive route that does not default to denial.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Test drift and versioning
Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.
Performance can decay when people, policy or measurement changes.
The audit names this model version, monitoring interval and stop threshold.
I’m stuck on this step
Reset: Re-read this authored instruction — “Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.” — and its success check, then attempt only this step.
Possible snag: A later cohort is scored without checking drift or this model version.
Correction: Compare base rates and calibration over time, then set a withdrawal threshold.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Map decision consequences
For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.
Prediction quality is not identical as decision value or legitimacy.
Worksheet separates forecast error from action harm.
I’m stuck on this step
Reset: Re-read this authored instruction — “For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.” — and its success check, then attempt only this step.
Possible snag: The result from “For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.” does not yet meet this declared check: Worksheet separates forecast error from action harm.
Correction: Return to the start of “Map decision consequences”, reduce complexity or pace, and repeat only the part needed to satisfy: “Worksheet separates forecast error from action harm.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Write evidence grade
Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.
A bounded grade prevents a narrow result becoming a destiny claim.
The conclusion includes context, uncertainty, groups, consequences and prohibited uses.
I’m stuck on this step
Reset: Re-read this authored instruction — “Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.” — and its success check, then attempt only this step.
Possible snag: The result from “Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.” does not yet meet this declared check: The conclusion includes context, uncertainty, groups, consequences and prohibited uses.
Correction: Return to the start of “Write evidence grade”, reduce complexity or pace, and repeat only the part needed to satisfy: “The conclusion includes context, uncertainty, groups, consequences and prohibited uses.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Compare it with identical-population base rate and error costs.
Call 78% impressive without a comparator.
Check observed frequencies in probability bins.
Treat a rank score as calibrated individual probability.
Mark transportability and fairness unresolved when a relevant group is absent.
Assume overall result applies equally to everyone.
Allow abstention and provide a human review plus correction route.
Force every case into a rank.
Method-structure checklist
10 of 10 structural checks present
- Ordered, Power-specific instructions — present
- Every activity has a success check — present
- Materials or supplied records are declared — present
- Measurement or assessment rule is present — present
- Tutorial-specific troubleshooting is present — present
- Stopping or escalation boundary is present — present
- Every activity has an adjacent alternative — present
- Correct-versus-incorrect comparison is present — present
- Evidence context is bound to the Power record — present
- Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.