Section 417 of 440
Complete canonical tutorial. This reader section contains the same teaching body as PWR-235 · Redundancy and resilience. Open the Power dossier.
PWR-235 · SELF GUIDED full tutorial
Map one critical function, expose shared backup failures and rehearse a safe fictional failover
This lesson teaches redundancy and resilience through a fictional service map. You will define the minimum function, trace dependencies, test whether a backup has capacity and independence, write a failover trigger, run a benign tabletop, test failure of both routes, recover and assign maintenance. A plan or spare object is not counted as resilience until the function test passes.
1 · Permission and limits
Know exactly what you may do
2 · Get ready
Gather what you need and check the starting conditions
What you need
- Fictional community information service with a primary laptop, spare laptop, one shared cloud account, paper telephone list, separate mailbox and four staff roles.
- Function/dependency table, common-cause map, failover trigger card and tabletop clock.
- Three supplied failure cards plus an observer rubric. F1 says the shared cloud account is locked while both laptops still work; the expected switch is to paper intake plus the separate mailbox. F2 says the trained coordinator is absent; the expected response is to name the authorised alternate role before proceeding. F3 says both paper forms and the separate mailbox are unavailable; minimum service cannot be met, so the tabletop stops and escalates to the service owner rather than inventing another route.
Before you start
- Label every item fictional and keep the exercise disconnected from live systems.
- Name the smallest acceptable service and fictional maximum interruption before mapping backups.
- Agree that any resemblance to a live incident ends the exercise and requires authorised review.
3 · The method
Follow these steps in order
- Define minimum function
State the fictional caller, verified callback output, acceptable quality and two-day maximum interruption.
Why: Resilience cannot be judged without a service target.
Check: The function is observable and deliberately smaller than “keep everything running.”
- Map primary dependencies
List people, building, power, device, account, network, supplier, information and decision authority.
Why: Invisible dependencies cause a backup to fail with the same upstream service.
Check: Every dependency has an owner or unknown marker.
- Describe backup capacity
Record what each alternative can deliver, for how long, to whom and with what loss of quality.
Why: A backup can exist but be too small or inaccessible.
Check: Capacity is compared directly with the minimum function.
- Test independence
Mark shared power, building, credential, supplier, data source and key person.
Why: Redundancy inside one failure domain is fragile.
Check: The map shows which failures take down both routes.
- Write the failover trigger
Specify the event, decision role, ordered switch steps, communication and time limit.
Why: Improvised switching can cause duplicate or conflicting work.
Check: A participant can follow the trigger without additional explanation.
- Run the benign failover
Draw a fictional primary-failure card, start the clock and switch records/roles to the approved alternative without touching live systems.
Why: A tabletop checks sequence and ownership safely.
Check: The minimum service resumes within the fictional target or the gap is recorded.
- Test both-routes failure
Draw the relevant common-cause or backup-failure card and choose degraded service, safe stop or external handoff.
Why: Resilience includes naming when fallback capacity no longer meets minimum service.
Check: No participant invents an unauthorised third route.
- Recover and reconcile
Return to the primary in the story, compare records, resolve duplicates and tell users which version is current.
Why: Failback can create a second incident.
Check: The current record and service owner are unambiguous.
- Assign maintenance
Set owners and dates for contacts, access, paper copies and the next test; verify one correction.
Why: Backups decay when contacts, credentials and role knowledge are never exercised.
Check: The repaired dependency passes a matched retest.
4 · Worked example
See the whole method used once
Scenario
A fictional service answers ordinary community questions within two working days using one cloud account; a paper telephone list and a separate shared mailbox are candidate fallbacks.
Walkthrough
- The learner defines minimum service as recording the caller’s question and giving a verified callback within two days.
- The dependency map shows both primary laptop and spare laptop use the same cloud account, so the spare is not independent for account failure.
- They choose the paper intake form plus separate mailbox as the tabletop route and switch in six minutes after an account-failure card.
- A second card removes the trained coordinator; the team discovers the alternate role is undocumented, assigns a role card and verifies it in a matched rerun.
Result
The tabletop exposes one false redundancy and repairs a key-person dependency. It does not validate the live organisation’s continuity.
5 · Right and wrong
Compare correct or safer execution with the common wrong version
| Moment | Right / safer | Wrong / riskier | Why it matters |
|---|---|---|---|
| Defining criticality | Choose the smallest service users actually need. | Mark every activity critical and leave no defensible recovery order. | If everything is critical, trade-offs and recovery order remain hidden. |
| Counting backups | Check fallback capacity and whether primary and backup share common failures. | Count a spare laptop on the same account as independent. | The shared credential can disable both. |
| Failing over | Use a trigger, owner and ordered communication. | Let several people improvise switches at once. | Uncoordinated action creates duplicate or lost work. |
| Ending capacity | Use degraded service, safe stop or authorised handoff. | Keep promising full service after both routes fail. | Honest limits prevent hidden unsafe work. |
6 · Common mistakes
Spot the error and apply the correction
| Mistake | Fix |
|---|---|
| Backup object with no capacity test | Compare volume, duration, access and quality against the minimum function. |
| Two resilience routes share the same account credentials. | Map credentials, identity provider and data source as dependencies. |
| Resilience depends on one undocumented role holder. | Document an alternate role and test cold handoff. |
| The primary route returns without reconciling fallback records. | Add record reconciliation, user notification and current-version checks. |
| A resilience test gap has no closure owner. | Assign owner, resources, due date and a matched verification card. |
7 · Practice
Turn the steps into a usable skill
First session
- Define minimum service and interruption target.
- Map primary and backup dependencies and common causes.
- Write the failover trigger and role handoff.
- Run one primary-failure and one both-routes-failure tabletop.
- Reconcile the fictional records, assign one correction and retest it.
Repeat plan
Review dependencies quarterly and after supplier, staff, building, technology or authority changes. The responsible organisation sets real exercise cadence; self-guided work remains fictional.
Progress when
- The learner distinguishes duplicate from independent capacity.
- The tabletop meets the declared switch target without role or version ambiguity.
- Both-routes failure ends in a planned degraded service, safe stop or authorised handoff.
Do not progress when
- The exercise touches a live system, real users, confidential data or regulated service.
- Recovery time, acceptable loss or authority cannot be set by the participants.
- A hidden fault injection or live outage is proposed.
8 · Check the result
Measure what changed
Tabletop continuity of one fictional critical function
How: Record minimum service, recovery target, capacity, shared dependencies, switch time, role handoff, both-routes response, failback and correction closure.
Good result: A good result resumes the minimum fictional service within target using a genuinely independent route and completes reconciliation.
This does not prove: It does not prove real capacity, equitable continuity, cybersecurity, compliance or recovery under a live disruption.
Self-check
- What is the smallest service and maximum interruption?
- Which single failure disables both primary and backup?
- Who declares failover and who owns the service afterward?
- How are records reconciled during recovery?
9 · Stop, adapt or get help
Keep the safety boundary practical
Stop and get help
- Stop if any step could affect a live service, account, employee, user or confidential record.
- Stop when participants lack authority to define service loss, failover or acceptable burden.
- Route real continuity, cybersecurity, clinical, emergency or infrastructure work to the accountable organisation and specialists.
Accessibility and adaptations
- Include accessible communications, alternate formats and support roles as dependencies and capacity requirements.
- Use readable role cards, diagrams, AAC and asynchronous tabletop input.
- Test the intended user’s access to the fallback; do not treat a technically working route as resilient if it excludes them.
10 · Evidence and limits
Why these instructions are here
- official guidance
NIST contingency-planning guidance describes business impact, alternate processing, testing, recovery and maintenance; a written plan alone is not demonstrated capacity.
Contingency Planning Guide for Federal Information Systems (NIST SP 800-34 Rev. 1) - official guidance
The Orange Book frames resilience within risk ownership, information, response and continual improvement rather than possession of spare components.
The Orange Book: Management of Risk — Principles and Concepts
Limits
- All practice uses a fictional service and tabletop cards.
- Redundancy is measured against one function, failure and capacity.
- A passed tabletop cannot validate live recovery time or outcome.
- A plan or backup object cannot guarantee resilience.
- Official standards specify practices rather than demonstrate outcomes; a written plan cannot establish capacity, recovery time or equitable continuity.
Open the complete canonical research register
- Official normative system supportLimiting / contraryContingency Planning Guide for Federal Information Systems (NIST SP 800-34 Rev. 1)
Marianne Swanson; Pauline Bowen; Amy Phillips; Dean Gallup; David Lynes; National Institute of Standards and Technology · 2010 · Official standard
- Official normative system supportLimiting / contraryThe Orange Book: Management of Risk — Principles and Concepts
HM Treasury · 2026 · Official standard
Read the complete evidence interpretation on the Power dossier.
Tutorial delivery controls
Learn, adapt, troubleshoot and resume
Progress is saved only in this browser on this device.
Step-by-step learner mode
Each activity includes its success check, a nearby accessible alternative and an “I’m stuck” correction path. Alternatives preserve the target where possible; when they change the task, Titan labels them as related rather than equivalent.
Define minimum function
State the fictional caller, verified callback output, acceptable quality and two-day maximum interruption.
Resilience cannot be judged without a service target.
The function is observable and deliberately smaller than “keep everything running.”
I’m stuck on this step
Reset: Re-read this authored instruction — “State the fictional caller, verified callback output, acceptable quality and two-day maximum interruption.” — and its success check, then attempt only this step.
Possible snag: The result from “State the fictional caller, verified callback output, acceptable quality and two-day maximum interruption.” does not yet meet this declared check: The function is observable and deliberately smaller than “keep everything running.”
Correction: Return to the start of “Define minimum function”, reduce complexity or pace, and repeat only the part needed to satisfy: “The function is observable and deliberately smaller than “keep everything running.””
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Map primary dependencies
List people, building, power, device, account, network, supplier, information and decision authority.
Invisible dependencies cause a backup to fail with the same upstream service.
Every dependency has an owner or unknown marker.
I’m stuck on this step
Reset: Re-read this authored instruction — “List people, building, power, device, account, network, supplier, information and decision authority.” — and its success check, then attempt only this step.
Possible snag: Two resilience routes share the same account credentials.
Correction: Map credentials, identity provider and data source as dependencies.
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Describe backup capacity
Record what each alternative can deliver, for how long, to whom and with what loss of quality.
A backup can exist but be too small or inaccessible.
Capacity is compared directly with the minimum function.
I’m stuck on this step
Reset: Re-read this authored instruction — “Record what each alternative can deliver, for how long, to whom and with what loss of quality.” — and its success check, then attempt only this step.
Possible snag: Backup object with no capacity test
Correction: Compare volume, duration, access and quality against the minimum function.
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Test independence
Mark shared power, building, credential, supplier, data source and key person.
Redundancy inside one failure domain is fragile.
The map shows which failures take down both routes.
I’m stuck on this step
Reset: Re-read this authored instruction — “Mark shared power, building, credential, supplier, data source and key person.” — and its success check, then attempt only this step.
Possible snag: The result from “Mark shared power, building, credential, supplier, data source and key person.” does not yet meet this declared check: The map shows which failures take down both routes.
Correction: Return to the start of “Test independence”, reduce complexity or pace, and repeat only the part needed to satisfy: “The map shows which failures take down both routes.”
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Write the failover trigger
Specify the event, decision role, ordered switch steps, communication and time limit.
Improvised switching can cause duplicate or conflicting work.
A participant can follow the trigger without additional explanation.
I’m stuck on this step
Reset: Re-read this authored instruction — “Specify the event, decision role, ordered switch steps, communication and time limit.” — and its success check, then attempt only this step.
Possible snag: The result from “Specify the event, decision role, ordered switch steps, communication and time limit.” does not yet meet this declared check: A participant can follow the trigger without additional explanation.
Correction: Return to the start of “Write the failover trigger”, reduce complexity or pace, and repeat only the part needed to satisfy: “A participant can follow the trigger without additional explanation.”
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Run the benign failover
Draw a fictional primary-failure card, start the clock and switch records/roles to the approved alternative without touching live systems.
A tabletop checks sequence and ownership safely.
The minimum service resumes within the fictional target or the gap is recorded.
I’m stuck on this step
Reset: Re-read this authored instruction — “Draw a fictional primary-failure card, start the clock and switch records/roles to the approved alternative without touching live systems.” — and its success check, then attempt only this step.
Possible snag: The result from “Draw a fictional primary-failure card, start the clock and switch records/roles to the approved alternative without touching live systems.” does not yet meet this declared check: The minimum service resumes within the fictional target or the gap is recorded.
Correction: Return to the start of “Run the benign failover”, reduce complexity or pace, and repeat only the part needed to satisfy: “The minimum service resumes within the fictional target or the gap is recorded.”
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Test both-routes failure
Draw the relevant common-cause or backup-failure card and choose degraded service, safe stop or external handoff.
Resilience includes naming when fallback capacity no longer meets minimum service.
No participant invents an unauthorised third route.
I’m stuck on this step
Reset: Re-read this authored instruction — “Draw the relevant common-cause or backup-failure card and choose degraded service, safe stop or external handoff.” — and its success check, then attempt only this step.
Possible snag: Resilience depends on one undocumented role holder.
Correction: Document an alternate role and test cold handoff.
Possible snag: A resilience test gap has no closure owner.
Correction: Assign owner, resources, due date and a matched verification card.
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Recover and reconcile
Return to the primary in the story, compare records, resolve duplicates and tell users which version is current.
Failback can create a second incident.
The current record and service owner are unambiguous.
I’m stuck on this step
Reset: Re-read this authored instruction — “Return to the primary in the story, compare records, resolve duplicates and tell users which version is current.” — and its success check, then attempt only this step.
Possible snag: The primary route returns without reconciling fallback records.
Correction: Add record reconciliation, user notification and current-version checks.
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Assign maintenance
Set owners and dates for contacts, access, paper copies and the next test; verify one correction.
Backups decay when contacts, credentials and role knowledge are never exercised.
The repaired dependency passes a matched retest.
I’m stuck on this step
Reset: Re-read this authored instruction — “Set owners and dates for contacts, access, paper copies and the next test; verify one correction.” — and its success check, then attempt only this step.
Possible snag: The result from “Set owners and dates for contacts, access, paper copies and the next test; verify one correction.” does not yet meet this declared check: The repaired dependency passes a matched retest.
Correction: Return to the start of “Assign maintenance”, reduce complexity or pace, and repeat only the part needed to satisfy: “The repaired dependency passes a matched retest.”
Stop / get help: Stop if any step could affect a live service, account, employee, user or confidential record.
Correct versus incorrect execution
These accessible process diagrams are built from the tutorial’s own right/wrong teaching. They are not anatomical illustrations and do not add technique beyond the canonical tutorial.
Choose the smallest service users actually need.
Mark every activity critical and leave no defensible recovery order.
Check fallback capacity and whether primary and backup share common failures.
Count a spare laptop on the same account as independent.
Use a trigger, owner and ordered communication.
Let several people improvise switches at once.
Use degraded service, safe stop or authorised handoff.
Keep promising full service after both routes fail.
Method-structure checklist
10 of 10 structural checks present
- Ordered, Power-specific instructions — present
- Every activity has a success check — present
- Materials or supplied records are declared — present
- Measurement or assessment rule is present — present
- Tutorial-specific troubleshooting is present — present
- Stopping or escalation boundary is present — present
- Every activity has an adjacent alternative — present
- Correct-versus-incorrect comparison is present — present
- Evidence context is bound to the Power record — present
- Planning metadata is present — present
The method-readiness band and presence checklist assess tutorial presentation and are separate from evidence quality for the underlying Power. They are automated editorial aids, not human approval.
Manual editorial sign-off: Pending. This tutorial must not display a human-approved state until an identified editor signs the exact content hash.