Revision 7T · Full Tutorial Edition · Updated 1 September 2026
PWR-221 · RESEARCH full tutorial
Audit a physiological digital-twin claim without treating a model as a person or a diagnosis
This research tutorial teaches how to inspect a physiological digital-twin study or product claim. You will define the clinical question, list represented and omitted variables, trace data and model versions, examine validation, calibration, subgroup errors and drift, and write a bounded verdict. No personal health data or treatment simulation is used.
What you will produceThe learner produces a one-page evidence audit that identifies the twin’s target, data, omissions, validation population, prediction error, uncertainty, update rules and prohibited decisions.
Method9 numbered Power-specific steps
Practice authorityComplete research method; no operational self-experiment
Full step-by-step individual tutorial · TLU-PWR-221
The learner produces a one-page evidence audit that identifies the twin’s target, data, omissions, validation population, prediction error, uncertainty, update rules and prohibited decisions.
Analyse a published paper, regulator document or fully fictional model card.
Recalculate simple reported error rates or calibration examples from supplied aggregate data.
Conclude that the evidence is insufficient, non-transportable or not decision-ready.
Qualified help is required for
Access to identifiable health data, model deployment, clinical validation or interpretation of a patient output.
Any diagnosis, treatment choice, triage, device use or prospective clinical study.
Cybersecurity, privacy, regulatory and change-control approval for a real system.
Never do this from the page alone
Enter personal health data into an unapproved model or upload it to a public tool.
Use a model output to diagnose, treat or advise a person.
Call a dashboard, simulation or confident prediction an exact digital copy of a body.
2 · Get ready
Gather what you need and check the starting conditions
What you need
The linked primary paper in the evidence notes, Development and Verification of a Digital Twin Patient Model to Predict Specific Treatment Response During the First 24 Hours of Sepsis, plus its available tables and supplement; no patient-level data or clinical decision is used.
An audit table with rows for question, inputs, omissions, version, population, comparator, calibration, uncertainty, subgroups, drift and decisions.
A calculator for supplied aggregate examples and a dated source log.
Before you start
Choose a non-personal source that exposes methods and validation details; do not use a live clinical dashboard.
Write the exact prediction target and time horizon before reading performance claims.
Set the verdict options: research signal, internally validated, externally validated for a named context, decision-impact evidence, or insufficient.
3 · The method
Follow these steps in order
Define the target question
Write the physiological outcome, population, time horizon and decision the model claims to inform.
Why: “Simulate the patient” is too broad to validate.
Check: The audit has one observable target and one declared non-use.
List what the twin represents
Separate measured inputs, inferred states, mechanistic equations, statistical components and outputs; list important body and context features omitted.
Why: A model is a selective representation, not a complete person.
Check: Every included and omitted element is explicit enough to challenge.
Trace data provenance
Record sites, dates, inclusion rules, missingness, coding, measurement timing, preprocessing and whether training and test people were separated.
Why: Leakage and timing errors can make a model appear more accurate than it is.
Check: The source, cohort and split method are named with unresolved gaps.
Lock the model version
Record software/model version, parameters, update frequency, human rules and any changes made after seeing results.
Why: A changing model cannot be interpreted from an old score without change control.
Check: The evaluated version and update rule are identifiable.
Examine validation
Identify internal, temporal and external validation; compare the model with a simple or existing comparator on the same cases.
Why: Performance in development data does not show transport to a new setting.
Check: The audit distinguishes each validation level and names the comparator.
Read error and calibration
Extract absolute errors, agreement or calibration across the clinically relevant range, not only a correlation or area-under-curve number.
Why: A model can rank cases while systematically over- or under-predicting risk.
Check: The audit states what the reported metric means and what it omits.
Check subgroups and failures
Look for missing-data behaviour, coding mistakes, extreme cases, subgroup errors, uncertainty, abstention and out-of-scope cases.
Why: Average performance can hide predictable harm.
Check: At least three failure modes and any untested subgroup are visible.
Test transport and drift claims
Ask what changes across site, population, treatment, sensor, coding and time and how the team detects performance drift.
Why: A validated model can become wrong after context changes.
Check: The audit names the revalidation triggers and whether they were tested.
Write the bounded verdict
State the strongest supported evidence level and list decisions the model must not make; separate prediction performance from patient benefit.
Why: Technical accuracy does not prove useful or safe care.
Check: The verdict includes uncertainty, governance gaps and the next evidence gate.
4 · Worked example
See the whole method used once
Scenario
A published sepsis digital-twin paper predicts one physiological response during the first 24 hours using historical hospital data; a promotional summary calls it “a virtual copy of every patient.”
Walkthrough
The learner rewrites the target as prediction of one named 24-hour response in the study cohort, not a whole patient.
They list coded observations and expert rules that are represented and note social context, future care and many physiological processes that are omitted.
They extract variable agreement and documented coding, rule and timing errors, then note that validation did not establish transport to another hospital.
They grade the claim “research signal with important implementation errors,” strike the virtual-copy wording and prohibit diagnosis or treatment use.
Result
The audit preserves the study’s narrow research contribution while rejecting whole-person, cross-setting and clinical-benefit claims.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
Right and wrong comparison
Moment
Right / safer
Wrong / riskier
Why it matters
Defining the twin
Name its variables, update rule, target and omissions.
Describe it as a complete digital patient.
Completeness cannot be tested and invites overreach.
Reading performance
Inspect calibration or absolute error and failure cases.
Use correlation or one headline score alone.
Association can hide systematic prediction error.
Testing transport
Require external validation in the intended population and setting.
Assume a model transfers because physiology is universal.
Care processes, coding, sensors and populations change the data-generating system.
Judging utility
Look for comparative decision and patient-outcome evidence.
Equate accurate prediction with better treatment.
A correct forecast can be ignored, misused or cause harm.
6 · Common mistakes
Spot the error and apply the correction
Common mistakes and corrections
Mistake
Fix
Auditing a digital twin without defining its prediction target.
Rewrite the model claim as one variable, horizon, population and decision before scoring evidence.
Accepting digital-twin validation with overlapping patient or time data.
Verify patient-level and time separation; mark validation uninterpretable if separation is unclear.
Reporting digital-twin accuracy without predicted-versus-observed calibration.
Extract predicted-versus-observed behaviour or state clearly that calibration is unreported.
Ignoring underrepresented groups in digital-twin performance review.
List clinically relevant groups and mark each tested, underpowered or absent.
Quoting digital-twin scores without model and data versions.
Attach every score to the exact model and data version; a changed model needs new evidence.
7 · Practice
Turn the steps into a usable skill
First session
Choose one paper and prewrite the target question and verdict scale.
Complete the representation, omission and provenance rows without reading promotional conclusions.
Extract validation, comparator, error, calibration and subgroup results.
Name three failure modes and one transport/drift trigger.
Write the bounded verdict and have another reader trace every sentence to the source.
Repeat plan
Audit one additional paper or regulator document each week for four weeks, alternating internally validated and externally evaluated examples. Progress by adding decision-impact evidence, not by handling personal data.
Progress when
The target, version, population and horizon are explicit before performance is interpreted.
Calibration, comparator, uncertainty, subgroups and failure cases are all addressed or marked unreported.
The verdict separates model accuracy, transport, clinical utility and regulatory authority.
Do not progress when
The source lacks methods, validation data or version information needed to audit the claim.
The exercise begins using identifiable data or advising a real person.
A promotional score is being treated as diagnosis, benefit or whole-person fidelity.
8 · Check the result
Measure what changed
Completeness and traceability of one digital-twin evidence audit
How: Score 12 required fields—target, horizon, population, inputs, omissions, provenance, version, comparator, calibration/error, subgroups, drift and prohibited use—as supported, absent or unclear, with a source location for each supported field.
Good result: All 12 fields are resolved or explicitly marked missing, every claim is traceable, and the verdict does not exceed the weakest critical field.
This does not prove: It does not validate the model, reproduce the study, assess a patient, prove clinical utility or confer regulatory clearance.
Self-check
Can you name three important parts of a person that this model does not represent?
Was validation internal, temporal or external, and were people separated across data splits?
What does the reported performance metric fail to show?
Which decisions remain prohibited even if the reported prediction is accurate?
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Stop if anyone proposes diagnosis, treatment, triage or device control from the audited output.
Route clinical, regulatory, privacy, cybersecurity or model-change questions to the accountable specialists.
Accessibility and adaptations
Use a plain-language model diagram and glossary for calibration, discrimination, missingness and drift.
Provide tables in screen-reader-compatible form and describe plots in words with the underlying numbers.
Allow a non-mathematical audit that marks a metric unexplained rather than forcing calculation without support.
10 · Evidence and limits
Why these instructions are here
primary research
A prospective digital-twin patient model showed variable agreement and implementation errors, illustrating why target, timing, code and external transport must be audited.
FDA’s computational-model credibility programme treats model credibility as context-of-use specific rather than evidence of a complete digital copy or universal clinical utility.
This is an evidence-audit tutorial, not model construction or clinical use.
A digital twin represents selected variables and cannot reproduce a whole person.
Even valid prediction does not establish patient benefit, fairness, transport or authority to act.
A limited physiological model cannot reproduce a whole person or authorize diagnosis and treatment.
The sepsis model showed variable agreement and common coding, expert-rule and timing errors; patient benefit and cross-setting transport were not established.
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
01
Define the target question
Write the physiological outcome, population, time horizon and decision the model claims to inform.
Why this step exists
“Simulate the patient” is too broad to validate.
Success check
The audit has one observable target and one declared non-use.
I’m stuck on this step
Reset: Re-read this authored instruction — “Write the physiological outcome, population, time horizon and decision the model claims to inform.” — and its success check, then attempt only this step.
Possible snag: Auditing a digital twin without defining its prediction target.
Correction: Rewrite the model claim as one variable, horizon, population and decision before scoring evidence.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
02
List what the twin represents
Separate measured inputs, inferred states, mechanistic equations, statistical components and outputs; list important body and context features omitted.
Why this step exists
A model is a selective representation, not a complete person.
Success check
Every included and omitted element is explicit enough to challenge.
I’m stuck on this step
Reset: Re-read this authored instruction — “Separate measured inputs, inferred states, mechanistic equations, statistical components and outputs; list important body and context features omitted.” — and its success check, then attempt only this step.
Possible snag: Ignoring underrepresented groups in digital-twin performance review.
Correction: List clinically relevant groups and mark each tested, underpowered or absent.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
03
Trace data provenance
Record sites, dates, inclusion rules, missingness, coding, measurement timing, preprocessing and whether training and test people were separated.
Why this step exists
Leakage and timing errors can make a model appear more accurate than it is.
Success check
The source, cohort and split method are named with unresolved gaps.
I’m stuck on this step
Reset: Re-read this authored instruction — “Record sites, dates, inclusion rules, missingness, coding, measurement timing, preprocessing and whether training and test people were separated.” — and its success check, then attempt only this step.
Possible snag: The result from “Record sites, dates, inclusion rules, missingness, coding, measurement timing, preprocessing and whether training and test people were separated.” does not yet meet this declared check: The source, cohort and split method are named with unresolved gaps.
Correction: Return to the start of “Trace data provenance”, reduce complexity or pace, and repeat only the part needed to satisfy: “The source, cohort and split method are named with unresolved gaps.”
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
04
Lock the model version
Record software/model version, parameters, update frequency, human rules and any changes made after seeing results.
Why this step exists
A changing model cannot be interpreted from an old score without change control.
Success check
The evaluated version and update rule are identifiable.
I’m stuck on this step
Reset: Re-read this authored instruction — “Record software/model version, parameters, update frequency, human rules and any changes made after seeing results.” — and its success check, then attempt only this step.
Possible snag: Quoting digital-twin scores without model and data versions.
Correction: Attach every score to the exact model and data version; a changed model needs new evidence.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
05
Examine validation
Identify internal, temporal and external validation; compare the model with a simple or existing comparator on the same cases.
Why this step exists
Performance in development data does not show transport to a new setting.
Success check
The audit distinguishes each validation level and names the comparator.
I’m stuck on this step
Reset: Re-read this authored instruction — “Identify internal, temporal and external validation; compare the model with a simple or existing comparator on the same cases.” — and its success check, then attempt only this step.
Possible snag: Accepting digital-twin validation with overlapping patient or time data.
Correction: Verify patient-level and time separation; mark validation uninterpretable if separation is unclear.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
06
Read error and calibration
Extract absolute errors, agreement or calibration across the clinically relevant range, not only a correlation or area-under-curve number.
Why this step exists
A model can rank cases while systematically over- or under-predicting risk.
Success check
The audit states what the reported metric means and what it omits.
I’m stuck on this step
Reset: Re-read this authored instruction — “Extract absolute errors, agreement or calibration across the clinically relevant range, not only a correlation or area-under-curve number.” — and its success check, then attempt only this step.
Possible snag: Reporting digital-twin accuracy without predicted-versus-observed calibration.
Correction: Extract predicted-versus-observed behaviour or state clearly that calibration is unreported.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
07
Check subgroups and failures
Look for missing-data behaviour, coding mistakes, extreme cases, subgroup errors, uncertainty, abstention and out-of-scope cases.
Why this step exists
Average performance can hide predictable harm.
Success check
At least three failure modes and any untested subgroup are visible.
I’m stuck on this step
Reset: Re-read this authored instruction — “Look for missing-data behaviour, coding mistakes, extreme cases, subgroup errors, uncertainty, abstention and out-of-scope cases.” — and its success check, then attempt only this step.
Possible snag: The result from “Look for missing-data behaviour, coding mistakes, extreme cases, subgroup errors, uncertainty, abstention and out-of-scope cases.” does not yet meet this declared check: At least three failure modes and any untested subgroup are visible.
Correction: Return to the start of “Check subgroups and failures”, reduce complexity or pace, and repeat only the part needed to satisfy: “At least three failure modes and any untested subgroup are visible.”
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
08
Test transport and drift claims
Ask what changes across site, population, treatment, sensor, coding and time and how the team detects performance drift.
Why this step exists
A validated model can become wrong after context changes.
Success check
The audit names the revalidation triggers and whether they were tested.
I’m stuck on this step
Reset: Re-read this authored instruction — “Ask what changes across site, population, treatment, sensor, coding and time and how the team detects performance drift.” — and its success check, then attempt only this step.
Possible snag: The result from “Ask what changes across site, population, treatment, sensor, coding and time and how the team detects performance drift.” does not yet meet this declared check: The audit names the revalidation triggers and whether they were tested.
Correction: Return to the start of “Test transport and drift claims”, reduce complexity or pace, and repeat only the part needed to satisfy: “The audit names the revalidation triggers and whether they were tested.”
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
09
Write the bounded verdict
State the strongest supported evidence level and list decisions the model must not make; separate prediction performance from patient benefit.
Why this step exists
Technical accuracy does not prove useful or safe care.
Success check
The verdict includes uncertainty, governance gaps and the next evidence gate.
I’m stuck on this step
Reset: Re-read this authored instruction — “State the strongest supported evidence level and list decisions the model must not make; separate prediction performance from patient benefit.” — and its success check, then attempt only this step.
Possible snag: The result from “State the strongest supported evidence level and list decisions the model must not make; separate prediction performance from patient benefit.” does not yet meet this declared check: The verdict includes uncertainty, governance gaps and the next evidence gate.
Correction: Return to the start of “Write the bounded verdict”, reduce complexity or pace, and repeat only the part needed to satisfy: “The verdict includes uncertainty, governance gaps and the next evidence gate.”
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Defining the twin — Completeness cannot be tested and invites overreach.
Correct / safer
Name its variables, update rule, target and omissions.
Wrong / riskier
Describe it as a complete digital patient.
Reading performance — Association can hide systematic prediction error.
Correct / safer
Inspect calibration or absolute error and failure cases.
Wrong / riskier
Use correlation or one headline score alone.
Testing transport — Care processes, coding, sensors and populations change the data-generating system.
Correct / safer
Require external validation in the intended population and setting.
Wrong / riskier
Assume a model transfers because physiology is universal.
Judging utility — A correct forecast can be ignored, misused or cause harm.
Correct / safer
Look for comparative decision and patient-outcome evidence.
Wrong / riskier
Equate accurate prediction with better treatment.
Method-structure checklist
10 of 10 structural checks present
✓ Ordered, Power-specific instructions — present
✓ Every activity has a success check — present
✓ Materials or supplied records are declared — present
✓ Measurement or assessment rule is present — present
✓ Tutorial-specific troubleshooting is present — present
✓ Stopping or escalation boundary is present — present
✓ Every activity has an adjacent alternative — present
✓ Correct-versus-incorrect comparison is present — present
✓ Evidence context is bound to the Power record — present
✓ Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.