# PWR-224 — Predictive capability modelling: learner worksheet

Release: Revision 7T · Full Tutorial Edition
Estimated time: Estimated 31 min reading and worksheet pass
Difficulty: Advanced
Equipment: Common household or practice equipment — A supplied fictional model report predicting whether 100 fictional trainees finish a six-week optional course. In the 2025 test set, 75 of 100 completed, the model classified 78 correctly and an always-complete base-rate rule classified 75 correctly.; Four supplied calibration rows: predicted 50%, n=20, 10 completed; predicted 70%, n=20, 14 completed; predicted 80%, n=20, 12 completed; predicted 90%, n=40, 39 completed. The accessibility subgroup has n=8 and its outcomes are suppressed. A 2026 drift snapshot has 60 completions among 100 people and no recalibration.; Audit worksheet, calculator and decision-consequence table.
Space: Desk / seated

Automated structural checklist: 10 of 10 structural checks present
Evidence quality/context: G5; Deep research depth
Automated method-quality band: Comprehensive
Method-rating basis: Structure coverage, instruction/check specificity, alternative distinctness, source troubleshooting, comparison depth and whether purpose/check fields are authored rather than derived.

Editorial review: Pending manual editorial sign off

## Before you begin
- [ ] I read the tutorial authority and stop conditions.
- [ ] I have the required equipment/space or a declared accessible alternative.

## Canonical source context

This research tutorial teaches a complete predictive-model audit. You will define one outcome with a forecast horizon; inspect target population and data splits; compare calibration with discrimination; check subgroups, uncertainty, drift, abstention and decision consequences; then write a bounded evidence grade. No personal ranking or consequential prediction is performed.

Target outcome: The learner produces a reproducible model-card audit that states what is predicted, for whom and when; reports validation plus subgroup limits; and rejects unsupported claims about worth, potential or destiny.

### Authority and setup conditions

- Confirm that every row is fictional or aggregate and no real person can be identified.
- Write the outcome, horizon, target population, allowed research question and prohibited uses.
- Predeclare the minimum evidence: separated test set, comparator, calibration, subgroup reporting, uncertainty and monitoring.

### Required or supplied materials

- A supplied fictional model report predicting whether 100 fictional trainees finish a six-week optional course. In the 2025 test set, 75 of 100 completed, the model classified 78 correctly and an always-complete base-rate rule classified 75 correctly.
- Four supplied calibration rows: predicted 50%, n=20, 10 completed; predicted 70%, n=20, 14 completed; predicted 80%, n=20, 12 completed; predicted 90%, n=40, 39 completed. The accessibility subgroup has n=8 and its outcomes are suppressed. A 2026 drift snapshot has 60 completions among 100 people and no recalibration.
- Audit worksheet, calculator and decision-consequence table.

## 1. Lock outcome and horizon

Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.

Why: Outcome blur allows a score to be reused for unrelated decisions.

Success check: The target can be marked observed or not observed at one declared time.

Accessible alternative: Complete the same research action — “Write exactly what event is predicted and by when, then separate that forecast from any judgement about a person.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “The target can be marked observed or not observed at one declared time.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 2. Define intended population

Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.

Why: Performance depends on who generated examples.

Success check: The audit names target population and every group to which this result must not transfer.

Accessible alternative: Complete the same research action — “Record the inclusion rules, exclusions, setting, base rate and groups missing from source data.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “The audit names target population and every group to which this result must not transfer.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 3. Inspect data separation

Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.

Why: Leakage can create impressive but unusable performance.

Success check: The final test is plausibly independent, or available evidence grade is stopped.

Accessible alternative: Complete the same research action — “Check whether people and time periods are separated across development, tuning and final test sets; also check whether outcome information leaked into the inputs.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “The final test is plausibly independent, or available evidence grade is stopped.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 4. Compare a simple baseline

Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.

Why: Extra complexity is useful only when it improves the declared forecast against a fair comparator.

Success check: Any claimed gain is shown for identical cases and outcome.

Accessible alternative: Complete the same research action — “Score this model against a transparent alternative, such as target population base rate or a one-variable rule, on identical test set.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “Any claimed gain is shown for identical cases and outcome.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 5. Separate discrimination from calibration

Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.

Why: A model can rank correctly while giving misleading probabilities.

Success check: Both ranking quality and probability accuracy are stated in plain English.

Accessible alternative: Complete the same research action — “Read ranking performance, then compare predicted probabilities with observed frequencies in each bin.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “Both ranking quality and probability accuracy are stated in plain English.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 6. Inspect subgroup and missing-data error

Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.

Why: An average can distribute errors unfairly or conceal exclusion.

Success check: The audit marks every subgroup as tested, underpowered or absent.

Accessible alternative: Complete the same research action — “Compare false positives, false negatives and calibration for each declared group, including records with missing inputs.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “The audit marks every subgroup as tested, underpowered or absent.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 7. Check uncertainty and abstention

Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.

Why: Forced predictions turn lack of knowledge into false precision.

Success check: Every uncertain case has a non-punitive route that does not default to denial.

Accessible alternative: Complete the same research action — “Find how this model signals out-of-scope or uncertain cases and what a human reviewer does when it abstains.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “Every uncertain case has a non-punitive route that does not default to denial.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 8. Test drift and versioning

Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.

Why: Performance can decay when people, policy or measurement changes.

Success check: The audit names this model version, monitoring interval and stop threshold.

Accessible alternative: Complete the same research action — “Compare later base rates and calibration with the original test period, then identify a trigger for revalidation or withdrawal.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “The audit names this model version, monitoring interval and stop threshold.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 9. Map decision consequences

For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.

Why: Prediction quality is not identical as decision value or legitimacy.

Success check: Worksheet separates forecast error from action harm.

Accessible alternative: Complete the same research action — “For false positives, false negatives, correct cases and abstentions, list who benefits, who is burdened, plus the appeal or correction required.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “Worksheet separates forecast error from action harm.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## 10. Write evidence grade

Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.

Why: A bounded grade prevents a narrow result becoming a destiny claim.

Success check: The conclusion includes context, uncertainty, groups, consequences and prohibited uses.

Accessible alternative: Complete the same research action — “Choose insufficient, internal research, external validation, decision-impact evidence or governed deployment evidence, then state the next gate.” — using speech-to-text, text-to-speech, enlarged text, keyboard-only navigation, shorter work blocks or a support person. Preserve this declared check: “The conclusion includes context, uncertainty, groups, consequences and prohibited uses.” Do not convert the research task into capability practice.

Alternative relationship: Target preserving when declared check is preserved

Learner notes:

________________________________________________________________________________

Completed: [ ]

## Reflection

What changed?

________________________________________________________________________________

What remains difficult?

________________________________________________________________________________

What will I repeat, adapt, ask for help with, or stop?

________________________________________________________________________________

---
Completion of this worksheet demonstrates tutorial participation only. It does not establish capability, qualification, safety clearance, diagnosis, treatment or independent validation.
