Revision 7T · Full Tutorial Edition · Updated 1 September 2026
PWR-224 · RESEARCH full tutorial
Evaluate a predictive capability model without turning a narrow forecast into a destiny score
This research tutorial teaches a complete predictive-model audit. You will define one outcome with a forecast horizon; inspect target population and data splits; compare calibration with discrimination; check subgroups, uncertainty, drift, abstention and decision consequences; then write a bounded evidence grade. No personal ranking or consequential prediction is performed.
What you will produceThe learner produces a reproducible model-card audit that states what is predicted, for whom and when; reports validation plus subgroup limits; and rejects unsupported claims about worth, potential or destiny.
Method10 numbered Power-specific steps
Practice authorityComplete research method; no operational self-experiment
Full step-by-step individual tutorial · TLU-PWR-224
The learner produces a reproducible model-card audit that states what is predicted, for whom and when; reports validation plus subgroup limits; and rejects unsupported claims about worth, potential or destiny.
Audit published aggregate model results or a fictional dataset.
Calculate simple calibration by comparing predicted and observed frequencies in supplied bins.
Recommend abstention, external validation or retirement when evidence is insufficient.
Qualified help is required for
Building or deploying a real model that affects education, employment, insurance, benefits, health, credit or liberty.
Access to personal data, fairness assessment for protected groups or impact evaluation.
Approval of thresholds, human override, appeals, monitoring and model changes in a real institution.
Never do this from the page alone
Score a real person’s potential, worth or future from this lesson.
Use a model trained for one outcome, population or horizon to decide another.
Hide subgroup error, uncertainty or abstention behind one overall accuracy number.
2 · Get ready
Gather what you need and check the starting conditions
What you need
A supplied fictional model report predicting whether 100 fictional trainees finish a six-week optional course. In the 2025 test set, 75 of 100 completed, the model classified 78 correctly and an always-complete base-rate rule classified 75 correctly.
Four supplied calibration rows: predicted 50%, n=20, 10 completed; predicted 70%, n=20, 14 completed; predicted 80%, n=20, 12 completed; predicted 90%, n=40, 39 completed. The accessibility subgroup has n=8 and its outcomes are suppressed. A 2026 drift snapshot has 60 completions among 100 people and no recalibration.
Audit worksheet, calculator and decision-consequence table.
Before you start
Confirm that every row is fictional or aggregate and no real person can be identified.
Write the outcome, horizon, target population, allowed research question and prohibited uses.
Predeclare the minimum evidence: separated test set, comparator, calibration, subgroup reporting, uncertainty and monitoring.
3 · The method
Follow these steps in order
Lock outcome and horizon
Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.
Why: Outcome blur allows a score to be reused for unrelated decisions.
Check: The target can be marked observed or not observed at one declared time.
Define intended population
Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.
Why: Performance depends on who generated examples.
Check: The audit names target population and every group to which this result must not transfer.
Inspect data separation
Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.
Why: Leakage can create impressive but unusable performance.
Check: The final test is plausibly independent, or available evidence grade is stopped.
Compare a simple baseline
Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.
Why: Extra complexity is useful only when it improves the declared forecast against a fair comparator.
Check: Any claimed gain is shown for identical cases and outcome.
Separate discrimination from calibration
Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.
Why: A model can rank correctly while giving misleading probabilities.
Check: Both ranking quality and probability accuracy are stated in plain English.
Inspect subgroup and missing-data error
Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.
Why: An average can distribute errors unfairly or conceal exclusion.
Check: The audit marks every subgroup as tested, underpowered or absent.
Check uncertainty and abstention
Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.
Why: Forced predictions turn lack of knowledge into false precision.
Check: Every uncertain case has a non-punitive route that does not default to denial.
Test drift and versioning
Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.
Why: Performance can decay when people, policy or measurement changes.
Check: The audit names this model version, monitoring interval and stop threshold.
Map decision consequences
For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.
Why: Prediction quality is not identical as decision value or legitimacy.
Check: Worksheet separates forecast error from action harm.
Write evidence grade
Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.
Why: A bounded grade prevents a narrow result becoming a destiny claim.
Check: The conclusion includes context, uncertainty, groups, consequences and prohibited uses.
4 · Worked example
See the whole method used once
Scenario
A fictional model claims to predict six-week course completion with 78% accuracy and is advertised as showing each trainee’s “future potential.”
Walkthrough
The learner rewrites the target outcome as completion of this optional course within six weeks in a fictional 2025 cohort.
They discover that 75% completed anyway, so this model’s three-point accuracy gain over the base-rate comparator is small.
The calibration table shows that only 60% of people in the 80% prediction bin completed; a small accessibility subgroup is not reported.
A 2026 snapshot has a different base rate. The learner grades this model for internal research only and prohibits admission, employment or personal-potential use.
Result
The audit preserves a narrow course-completion signal but rejects calibrated-probability, subgroup, drift, benefit and destiny claims.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
Right and wrong comparison
Moment
Right / safer
Wrong / riskier
Why it matters
Reading accuracy
Compare it with identical-population base rate and error costs.
Call 78% impressive without a comparator.
A high number may barely beat a trivial rule.
Reading probability
Check observed frequencies in probability bins.
Treat a rank score as calibrated individual probability.
Discrimination and calibration answer different questions; one cannot stand in for the other.
Handling missing groups
Mark transportability and fairness unresolved when a relevant group is absent.
Assume overall result applies equally to everyone.
Unreported groups cannot inherit average validity.
Using uncertainty
Allow abstention and provide a human review plus correction route.
Force every case into a rank.
Forced certainty magnifies error when a case falls outside available evidence.
6 · Common mistakes
Spot the error and apply the correction
Common mistakes and corrections
Mistake
Fix
The course-completion forecast is relabelled as a measure of future potential.
Rewrite this result with one outcome, horizon, population and model version every time it is cited.
The report quotes accuracy without a baseline comparator.
Add a transparent base-rate or existing-process comparison on identical test cases.
The audit treats a ranking score as an individual probability.
Build predicted-versus-observed bins or mark probability claims unsupported.
Only favourable subgroup results are reported after looking at source data.
Predeclare relevant groups and report all of them, including small or missing samples, without cherry-picking.
A later cohort is scored without checking drift or this model version.
Compare base rates and calibration over time, then set a withdrawal threshold.
7 · Practice
Turn the steps into a usable skill
First session
Predeclare the fictional outcome, population, evidence gate and prohibited uses.
Check the splits, leakage and base-rate comparator.
Calculate the supplied calibration bins plus error types.
Audit the subgroup, missing-data, uncertainty and drift fields.
Map the decision consequences, then write a bounded evidence grade.
Repeat plan
Audit one new fictional or published aggregate model each week for four weeks. Alternate a well-calibrated example with a poorly calibrated one. Progress only to decision-impact studies, never to personal scoring.
Progress when
Always state target outcome, horizon, population, version and comparator before quoting a score.
Calibration, subgroup errors, uncertainty and drift are interpreted correctly.
The final grade stays within the available validation and decision-impact evidence.
Do not progress when
Data separation, target definition or model version cannot be established.
Exercise shifts to a real person or a consequential decision.
A missing subgroup or drift failure is being treated as proof of no problem.
8 · Check the result
Measure what changed
Completeness and correctness of a predictive-model evidence audit
How: Score ten fields—outcome/horizon, population, split, comparator, discrimination, calibration, subgroups, uncertainty, drift and consequences—with traceable values plus one evidence grade.
Good result: All fields are correct or explicitly unresolved, this grade matches the weakest essential field and no personal-potential claim remains.
This does not prove: It does not validate a real model, establish causal benefit, prove fairness or authorise any consequential decision.
Self-check
What trivial comparator must this model beat?
Can you explain the difference between discrimination and calibration?
Which group or context is absent, what claim must therefore stop?
What happens to a person when this model abstains or makes each error type?
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
Stop if real personal or protected data is introduced without authorised governance.
Stop if output is proposed for admission, employment, insurance, credit, health, benefits, discipline or another consequential decision.
Route model deployment, legal, privacy, fairness, accessibility and appeal design to accountable specialists and affected people.
Accessibility and adaptations
Explain accuracy, calibration and false-positive trade-offs with frequencies and plain-language tables rather than formulas alone.
Provide screen-reader-friendly tables plus text descriptions of reliability plots.
Allow calculators and worked examples; do not equate slower arithmetic with weaker model judgement.
10 · Evidence and limits
Why these instructions are here
official guidance
NIST’s AI Risk Management Framework calls for context, validity, reliability, transparency and risk governance rather than treating a model score as universal truth.
A prospective digital-twin prediction study reported variable agreement and implementation errors, illustrating the limits of transporting model outputs across contexts.
This tutorial evaluates evidence; it does not build, tune or deploy a model.
A prediction is tied to one outcome, horizon, population, dataset and version.
No model score measures human worth, destiny or total capability.
A context-specific forecast cannot measure worth, destiny or total capability.
The prospective digital-twin study showed variable agreement and common implementation errors; no cross-context universal capability prediction or improved decision outcome was established.
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
01
Lock outcome and horizon
Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.
Why this step exists
Outcome blur allows a score to be reused for unrelated decisions.
Success check
The target can be marked observed or not observed at one declared time.
I’m stuck on this step
Reset: Re-read this authored instruction — “Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.” — and its success check, then attempt only this step.
Possible snag: The course-completion forecast is relabelled as a measure of future potential.
Correction: Rewrite this result with one outcome, horizon, population and model version every time it is cited.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
02
Define intended population
Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.
Why this step exists
Performance depends on who generated examples.
Success check
The audit names target population and every group to which this result must not transfer.
I’m stuck on this step
Reset: Re-read this authored instruction — “Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.” — and its success check, then attempt only this step.
Possible snag: Only favourable subgroup results are reported after looking at source data.
Correction: Predeclare relevant groups and report all of them, including small or missing samples, without cherry-picking.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
03
Inspect data separation
Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.
Why this step exists
Leakage can create impressive but unusable performance.
Success check
The final test is plausibly independent, or available evidence grade is stopped.
I’m stuck on this step
Reset: Re-read this authored instruction — “Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.” — and its success check, then attempt only this step.
Possible snag: The result from “Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.” does not yet meet this declared check: The final test is plausibly independent, or available evidence grade is stopped.
Correction: Return to the start of “Inspect data separation”, reduce complexity or pace, and repeat only the part needed to satisfy: “The final test is plausibly independent, or available evidence grade is stopped.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
04
Compare a simple baseline
Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.
Why this step exists
Extra complexity is useful only when it improves the declared forecast against a fair comparator.
Success check
Any claimed gain is shown for identical cases and outcome.
I’m stuck on this step
Reset: Re-read this authored instruction — “Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.” — and its success check, then attempt only this step.
Possible snag: The report quotes accuracy without a baseline comparator.
Correction: Add a transparent base-rate or existing-process comparison on identical test cases.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
05
Separate discrimination from calibration
Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.
Why this step exists
A model can rank correctly while giving misleading probabilities.
Success check
Both ranking quality and probability accuracy are stated in plain English.
I’m stuck on this step
Reset: Re-read this authored instruction — “Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.” — and its success check, then attempt only this step.
Possible snag: The audit treats a ranking score as an individual probability.
Correction: Build predicted-versus-observed bins or mark probability claims unsupported.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
06
Inspect subgroup and missing-data error
Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.
Why this step exists
An average can distribute errors unfairly or conceal exclusion.
Success check
The audit marks every subgroup as tested, underpowered or absent.
I’m stuck on this step
Reset: Re-read this authored instruction — “Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.” — and its success check, then attempt only this step.
Possible snag: The result from “Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.” does not yet meet this declared check: The audit marks every subgroup as tested, underpowered or absent.
Correction: Return to the start of “Inspect subgroup and missing-data error”, reduce complexity or pace, and repeat only the part needed to satisfy: “The audit marks every subgroup as tested, underpowered or absent.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
07
Check uncertainty and abstention
Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.
Why this step exists
Forced predictions turn lack of knowledge into false precision.
Success check
Every uncertain case has a non-punitive route that does not default to denial.
I’m stuck on this step
Reset: Re-read this authored instruction — “Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.” — and its success check, then attempt only this step.
Possible snag: The result from “Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.” does not yet meet this declared check: Every uncertain case has a non-punitive route that does not default to denial.
Correction: Return to the start of “Check uncertainty and abstention”, reduce complexity or pace, and repeat only the part needed to satisfy: “Every uncertain case has a non-punitive route that does not default to denial.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
08
Test drift and versioning
Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.
Why this step exists
Performance can decay when people, policy or measurement changes.
Success check
The audit names this model version, monitoring interval and stop threshold.
I’m stuck on this step
Reset: Re-read this authored instruction — “Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.” — and its success check, then attempt only this step.
Possible snag: A later cohort is scored without checking drift or this model version.
Correction: Compare base rates and calibration over time, then set a withdrawal threshold.
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
09
Map decision consequences
For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.
Why this step exists
Prediction quality is not identical as decision value or legitimacy.
Success check
Worksheet separates forecast error from action harm.
I’m stuck on this step
Reset: Re-read this authored instruction — “For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.” — and its success check, then attempt only this step.
Possible snag: The result from “For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.” does not yet meet this declared check: Worksheet separates forecast error from action harm.
Correction: Return to the start of “Map decision consequences”, reduce complexity or pace, and repeat only the part needed to satisfy: “Worksheet separates forecast error from action harm.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
10
Write evidence grade
Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.
Why this step exists
A bounded grade prevents a narrow result becoming a destiny claim.
Success check
The conclusion includes context, uncertainty, groups, consequences and prohibited uses.
I’m stuck on this step
Reset: Re-read this authored instruction — “Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.” — and its success check, then attempt only this step.
Possible snag: The result from “Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.” does not yet meet this declared check: The conclusion includes context, uncertainty, groups, consequences and prohibited uses.
Correction: Return to the start of “Write evidence grade”, reduce complexity or pace, and repeat only the part needed to satisfy: “The conclusion includes context, uncertainty, groups, consequences and prohibited uses.”
Stop / get help: Stop if real personal or protected data is introduced without authorised governance.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Reading accuracy — A high number may barely beat a trivial rule.
Correct / safer
Compare it with identical-population base rate and error costs.
Wrong / riskier
Call 78% impressive without a comparator.
Reading probability — Discrimination and calibration answer different questions; one cannot stand in for the other.
Correct / safer
Check observed frequencies in probability bins.
Wrong / riskier
Treat a rank score as calibrated individual probability.
Handling missing groups — Unreported groups cannot inherit average validity.
Correct / safer
Mark transportability and fairness unresolved when a relevant group is absent.
Wrong / riskier
Assume overall result applies equally to everyone.
Using uncertainty — Forced certainty magnifies error when a case falls outside available evidence.
Correct / safer
Allow abstention and provide a human review plus correction route.
Wrong / riskier
Force every case into a rank.
Method-structure checklist
10 of 10 structural checks present
✓ Ordered, Power-specific instructions — present
✓ Every activity has a success check — present
✓ Materials or supplied records are declared — present
✓ Measurement or assessment rule is present — present
✓ Tutorial-specific troubleshooting is present — present
✓ Stopping or escalation boundary is present — present
✓ Every activity has an adjacent alternative — present
✓ Correct-versus-incorrect comparison is present — present
✓ Evidence context is bound to the Power record — present
✓ Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.