Section 403 of 440
Complete canonical tutorial. This reader section contains the same teaching body as PWR-221 · Physiological digital twin. Open the Power dossier.
PWR-221 · RESEARCH full tutorial
Audit a physiological digital-twin claim without treating a model as a person or a diagnosis
This research tutorial teaches how to inspect a physiological digital-twin study or product claim. You will define the clinical question, list represented and omitted variables, trace data and model versions, examine validation, calibration, subgroup errors and drift, and write a bounded verdict. No personal health data or treatment simulation is used.
1 · Permission and limits
Know exactly what you may do
2 · Get ready
Gather what you need and check the starting conditions
What you need
- The linked primary paper in the evidence notes, Development and Verification of a Digital Twin Patient Model to Predict Specific Treatment Response During the First 24 Hours of Sepsis, plus its available tables and supplement; no patient-level data or clinical decision is used.
- An audit table with rows for question, inputs, omissions, version, population, comparator, calibration, uncertainty, subgroups, drift and decisions.
- A calculator for supplied aggregate examples and a dated source log.
Before you start
- Choose a non-personal source that exposes methods and validation details; do not use a live clinical dashboard.
- Write the exact prediction target and time horizon before reading performance claims.
- Set the verdict options: research signal, internally validated, externally validated for a named context, decision-impact evidence, or insufficient.
3 · The method
Follow these steps in order
- Define the target question
Write the physiological outcome, population, time horizon and decision the model claims to inform.
Why: “Simulate the patient” is too broad to validate.
Check: The audit has one observable target and one declared non-use.
- List what the twin represents
Separate measured inputs, inferred states, mechanistic equations, statistical components and outputs; list important body and context features omitted.
Why: A model is a selective representation, not a complete person.
Check: Every included and omitted element is explicit enough to challenge.
- Trace data provenance
Record sites, dates, inclusion rules, missingness, coding, measurement timing, preprocessing and whether training and test people were separated.
Why: Leakage and timing errors can make a model appear more accurate than it is.
Check: The source, cohort and split method are named with unresolved gaps.
- Lock the model version
Record software/model version, parameters, update frequency, human rules and any changes made after seeing results.
Why: A changing model cannot be interpreted from an old score without change control.
Check: The evaluated version and update rule are identifiable.
- Examine validation
Identify internal, temporal and external validation; compare the model with a simple or existing comparator on the same cases.
Why: Performance in development data does not show transport to a new setting.
Check: The audit distinguishes each validation level and names the comparator.
- Read error and calibration
Extract absolute errors, agreement or calibration across the clinically relevant range, not only a correlation or area-under-curve number.
Why: A model can rank cases while systematically over- or under-predicting risk.
Check: The audit states what the reported metric means and what it omits.
- Check subgroups and failures
Look for missing-data behaviour, coding mistakes, extreme cases, subgroup errors, uncertainty, abstention and out-of-scope cases.
Why: Average performance can hide predictable harm.
Check: At least three failure modes and any untested subgroup are visible.
- Test transport and drift claims
Ask what changes across site, population, treatment, sensor, coding and time and how the team detects performance drift.
Why: A validated model can become wrong after context changes.
Check: The audit names the revalidation triggers and whether they were tested.
- Write the bounded verdict
State the strongest supported evidence level and list decisions the model must not make; separate prediction performance from patient benefit.
Why: Technical accuracy does not prove useful or safe care.
Check: The verdict includes uncertainty, governance gaps and the next evidence gate.
4 · Worked example
See the whole method used once
Scenario
A published sepsis digital-twin paper predicts one physiological response during the first 24 hours using historical hospital data; a promotional summary calls it “a virtual copy of every patient.”
Walkthrough
- The learner rewrites the target as prediction of one named 24-hour response in the study cohort, not a whole patient.
- They list coded observations and expert rules that are represented and note social context, future care and many physiological processes that are omitted.
- They extract variable agreement and documented coding, rule and timing errors, then note that validation did not establish transport to another hospital.
- They grade the claim “research signal with important implementation errors,” strike the virtual-copy wording and prohibit diagnosis or treatment use.
Result
The audit preserves the study’s narrow research contribution while rejecting whole-person, cross-setting and clinical-benefit claims.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
| Moment | Right / safer | Wrong / riskier | Why it matters |
|---|---|---|---|
| Defining the twin | Name its variables, update rule, target and omissions. | Describe it as a complete digital patient. | Completeness cannot be tested and invites overreach. |
| Reading performance | Inspect calibration or absolute error and failure cases. | Use correlation or one headline score alone. | Association can hide systematic prediction error. |
| Testing transport | Require external validation in the intended population and setting. | Assume a model transfers because physiology is universal. | Care processes, coding, sensors and populations change the data-generating system. |
| Judging utility | Look for comparative decision and patient-outcome evidence. | Equate accurate prediction with better treatment. | A correct forecast can be ignored, misused or cause harm. |
6 · Common mistakes
Spot the error and apply the correction
| Mistake | Fix |
|---|---|
| Auditing a digital twin without defining its prediction target. | Rewrite the model claim as one variable, horizon, population and decision before scoring evidence. |
| Accepting digital-twin validation with overlapping patient or time data. | Verify patient-level and time separation; mark validation uninterpretable if separation is unclear. |
| Reporting digital-twin accuracy without predicted-versus-observed calibration. | Extract predicted-versus-observed behaviour or state clearly that calibration is unreported. |
| Ignoring underrepresented groups in digital-twin performance review. | List clinically relevant groups and mark each tested, underpowered or absent. |
| Quoting digital-twin scores without model and data versions. | Attach every score to the exact model and data version; a changed model needs new evidence. |
7 · Practice
Turn the steps into a usable skill
First session
- Choose one paper and prewrite the target question and verdict scale.
- Complete the representation, omission and provenance rows without reading promotional conclusions.
- Extract validation, comparator, error, calibration and subgroup results.
- Name three failure modes and one transport/drift trigger.
- Write the bounded verdict and have another reader trace every sentence to the source.
Repeat plan
Audit one additional paper or regulator document each week for four weeks, alternating internally validated and externally evaluated examples. Progress by adding decision-impact evidence, not by handling personal data.
Progress when
- The target, version, population and horizon are explicit before performance is interpreted.
- Calibration, comparator, uncertainty, subgroups and failure cases are all addressed or marked unreported.
- The verdict separates model accuracy, transport, clinical utility and regulatory authority.
Do not progress when
- The source lacks methods, validation data or version information needed to audit the claim.
- The exercise begins using identifiable data or advising a real person.
- A promotional score is being treated as diagnosis, benefit or whole-person fidelity.
8 · Check the result
Measure what changed
Completeness and traceability of one digital-twin evidence audit
How: Score 12 required fields—target, horizon, population, inputs, omissions, provenance, version, comparator, calibration/error, subgroups, drift and prohibited use—as supported, absent or unclear, with a source location for each supported field.
Good result: All 12 fields are resolved or explicitly marked missing, every claim is traceable, and the verdict does not exceed the weakest critical field.
This does not prove: It does not validate the model, reproduce the study, assess a patient, prove clinical utility or confer regulatory clearance.
Self-check
- Can you name three important parts of a person that this model does not represent?
- Was validation internal, temporal or external, and were people separated across data splits?
- What does the reported performance metric fail to show?
- Which decisions remain prohibited even if the reported prediction is accurate?
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
- Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
- Stop if anyone proposes diagnosis, treatment, triage or device control from the audited output.
- Route clinical, regulatory, privacy, cybersecurity or model-change questions to the accountable specialists.
Accessibility and adaptations
- Use a plain-language model diagram and glossary for calibration, discrimination, missingness and drift.
- Provide tables in screen-reader-compatible form and describe plots in words with the underlying numbers.
- Allow a non-mathematical audit that marks a metric unexplained rather than forcing calculation without support.
10 · Evidence and limits
Why these instructions are here
- primary research
A prospective digital-twin patient model showed variable agreement and implementation errors, illustrating why target, timing, code and external transport must be audited.
Development and Verification of a Digital Twin Patient Model to Predict Specific Treatment Response During the First 24 Hours of Sepsis - official guidance
FDA’s computational-model credibility programme treats model credibility as context-of-use specific rather than evidence of a complete digital copy or universal clinical utility.
Credibility of Computational Models Program: Research on Computational Models and Simulation Associated with Medical Devices
Limits
- This is an evidence-audit tutorial, not model construction or clinical use.
- A digital twin represents selected variables and cannot reproduce a whole person.
- Even valid prediction does not establish patient benefit, fairness, transport or authority to act.
- A limited physiological model cannot reproduce a whole person or authorize diagnosis and treatment.
- The sepsis model showed variable agreement and common coding, expert-rule and timing errors; patient benefit and cross-setting transport were not established.
Open the complete canonical research register
- Primary empirical supportLimiting / contraryDevelopment and Verification of a Digital Twin Patient Model to Predict Specific Treatment Response During the First 24 Hours of Sepsis
A Lal; G Li; E Cubro; S Chalmers; H Li; V Herasevich; Y Dong; B W Pickering; O Kilickaya; O Gajic · 2020 · Primary research
- Limiting / contraryOfficial boundary contextArtificial Intelligence Risk Management Framework (AI RMF 1.0)
Elham Tabassi; National Institute of Standards and Technology · 2023 · Official standard
- Limiting / contraryOfficial boundary contextNIST Privacy Framework: A Tool for Improving Privacy Through Enterprise Risk Management, Version 1.0
National Institute of Standards and Technology · 2020 · Official standard
- Limiting / contraryOfficial boundary contextCredibility of Computational Models Program: Research on Computational Models and Simulation Associated with Medical Devices
U.S. Food and Drug Administration, Office of Science and Engineering Laboratories · 2023 · Official guidance
- Limiting / contraryOfficial boundary contextClinical Decision Support Software
U.S. Food and Drug Administration · 2026 · Official guidance
- Limiting / contraryOfficial boundary contextCybersecurity in Medical Devices: Quality Management System Considerations and Content of Premarket Submissions
United States Food and Drug Administration · 2026 · Official guidance
- Limiting / contraryOfficial boundary contextMarketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions
U.S. Food and Drug Administration · 2025 · Official guidance
Read the complete evidence interpretation on the Power dossier.
Tutorial delivery controls
Learn, adapt, troubleshoot and resume
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
Define the target question
Write the physiological outcome, population, time horizon and decision the model claims to inform.
“Simulate the patient” is too broad to validate.
The audit has one observable target and one declared non-use.
I’m stuck on this step
Reset: Re-read this authored instruction — “Write the physiological outcome, population, time horizon and decision the model claims to inform.” — and its success check, then attempt only this step.
Possible snag: Auditing a digital twin without defining its prediction target.
Correction: Rewrite the model claim as one variable, horizon, population and decision before scoring evidence.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
List what the twin represents
Separate measured inputs, inferred states, mechanistic equations, statistical components and outputs; list important body and context features omitted.
A model is a selective representation, not a complete person.
Every included and omitted element is explicit enough to challenge.
I’m stuck on this step
Reset: Re-read this authored instruction — “Separate measured inputs, inferred states, mechanistic equations, statistical components and outputs; list important body and context features omitted.” — and its success check, then attempt only this step.
Possible snag: Ignoring underrepresented groups in digital-twin performance review.
Correction: List clinically relevant groups and mark each tested, underpowered or absent.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Trace data provenance
Record sites, dates, inclusion rules, missingness, coding, measurement timing, preprocessing and whether training and test people were separated.
Leakage and timing errors can make a model appear more accurate than it is.
The source, cohort and split method are named with unresolved gaps.
I’m stuck on this step
Reset: Re-read this authored instruction — “Record sites, dates, inclusion rules, missingness, coding, measurement timing, preprocessing and whether training and test people were separated.” — and its success check, then attempt only this step.
Possible snag: The result from “Record sites, dates, inclusion rules, missingness, coding, measurement timing, preprocessing and whether training and test people were separated.” does not yet meet this declared check: The source, cohort and split method are named with unresolved gaps.
Correction: Return to the start of “Trace data provenance”, reduce complexity or pace, and repeat only the part needed to satisfy: “The source, cohort and split method are named with unresolved gaps.”
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Lock the model version
Record software/model version, parameters, update frequency, human rules and any changes made after seeing results.
A changing model cannot be interpreted from an old score without change control.
The evaluated version and update rule are identifiable.
I’m stuck on this step
Reset: Re-read this authored instruction — “Record software/model version, parameters, update frequency, human rules and any changes made after seeing results.” — and its success check, then attempt only this step.
Possible snag: Quoting digital-twin scores without model and data versions.
Correction: Attach every score to the exact model and data version; a changed model needs new evidence.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Examine validation
Identify internal, temporal and external validation; compare the model with a simple or existing comparator on the same cases.
Performance in development data does not show transport to a new setting.
The audit distinguishes each validation level and names the comparator.
I’m stuck on this step
Reset: Re-read this authored instruction — “Identify internal, temporal and external validation; compare the model with a simple or existing comparator on the same cases.” — and its success check, then attempt only this step.
Possible snag: Accepting digital-twin validation with overlapping patient or time data.
Correction: Verify patient-level and time separation; mark validation uninterpretable if separation is unclear.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Read error and calibration
Extract absolute errors, agreement or calibration across the clinically relevant range, not only a correlation or area-under-curve number.
A model can rank cases while systematically over- or under-predicting risk.
The audit states what the reported metric means and what it omits.
I’m stuck on this step
Reset: Re-read this authored instruction — “Extract absolute errors, agreement or calibration across the clinically relevant range, not only a correlation or area-under-curve number.” — and its success check, then attempt only this step.
Possible snag: Reporting digital-twin accuracy without predicted-versus-observed calibration.
Correction: Extract predicted-versus-observed behaviour or state clearly that calibration is unreported.
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Check subgroups and failures
Look for missing-data behaviour, coding mistakes, extreme cases, subgroup errors, uncertainty, abstention and out-of-scope cases.
Average performance can hide predictable harm.
At least three failure modes and any untested subgroup are visible.
I’m stuck on this step
Reset: Re-read this authored instruction — “Look for missing-data behaviour, coding mistakes, extreme cases, subgroup errors, uncertainty, abstention and out-of-scope cases.” — and its success check, then attempt only this step.
Possible snag: The result from “Look for missing-data behaviour, coding mistakes, extreme cases, subgroup errors, uncertainty, abstention and out-of-scope cases.” does not yet meet this declared check: At least three failure modes and any untested subgroup are visible.
Correction: Return to the start of “Check subgroups and failures”, reduce complexity or pace, and repeat only the part needed to satisfy: “At least three failure modes and any untested subgroup are visible.”
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Test transport and drift claims
Ask what changes across site, population, treatment, sensor, coding and time and how the team detects performance drift.
A validated model can become wrong after context changes.
The audit names the revalidation triggers and whether they were tested.
I’m stuck on this step
Reset: Re-read this authored instruction — “Ask what changes across site, population, treatment, sensor, coding and time and how the team detects performance drift.” — and its success check, then attempt only this step.
Possible snag: The result from “Ask what changes across site, population, treatment, sensor, coding and time and how the team detects performance drift.” does not yet meet this declared check: The audit names the revalidation triggers and whether they were tested.
Correction: Return to the start of “Test transport and drift claims”, reduce complexity or pace, and repeat only the part needed to satisfy: “The audit names the revalidation triggers and whether they were tested.”
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Write the bounded verdict
State the strongest supported evidence level and list decisions the model must not make; separate prediction performance from patient benefit.
Technical accuracy does not prove useful or safe care.
The verdict includes uncertainty, governance gaps and the next evidence gate.
I’m stuck on this step
Reset: Re-read this authored instruction — “State the strongest supported evidence level and list decisions the model must not make; separate prediction performance from patient benefit.” — and its success check, then attempt only this step.
Possible snag: The result from “State the strongest supported evidence level and list decisions the model must not make; separate prediction performance from patient benefit.” does not yet meet this declared check: The verdict includes uncertainty, governance gaps and the next evidence gate.
Correction: Return to the start of “Write the bounded verdict”, reduce complexity or pace, and repeat only the part needed to satisfy: “The verdict includes uncertainty, governance gaps and the next evidence gate.”
Stop / get help: Stop if identifiable or sensitive health data enters the exercise without the required approved system and authority.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Name its variables, update rule, target and omissions.
Describe it as a complete digital patient.
Inspect calibration or absolute error and failure cases.
Use correlation or one headline score alone.
Require external validation in the intended population and setting.
Assume a model transfers because physiology is universal.
Look for comparative decision and patient-outcome evidence.
Equate accurate prediction with better treatment.
Method-structure checklist
10 of 10 structural checks present
- Ordered, Power-specific instructions — present
- Every activity has a success check — present
- Materials or supplied records are declared — present
- Measurement or assessment rule is present — present
- Tutorial-specific troubleshooting is present — present
- Stopping or escalation boundary is present — present
- Every activity has an adjacent alternative — present
- Correct-versus-incorrect comparison is present — present
- Evidence context is bound to the Power record — present
- Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.