Give Each Dependency Failure a Row and Four Questions
Offline Failure Mode Worksheet
Companion to Medical Device Connectivity · Updated 2026-09-29
Use this for every device with any network dependency: when you define its degraded modes, and again whenever a feature adds a connection to anything outside the device. Systems, software, and human-factors engineers fill it in, and clinical or process users review the last column; “Network Failure Is a Normal Operating Condition” works it on the smart infusion pump. The rows are the same for a patient-controlled analgesia pump, an anesthesia workstation, a smart bed that forwards bed-exit alarms, or a dialysis machine; the answers differ.
Fill in each column as follows:
Degraded mode:: The degraded mode the failure puts the device in: one of the four named in “A Device’s Safety Should Never Depend on the Network” (Isolated, Local sign-in, Unverified time, Reduced service), or a mode of the device’s own if none fits. What continues:: The functions that keep working, starting with the essential function. In the identity-provider and expired-credential rows, list how therapy changes, alarm acknowledgment, and stopping a therapy continue without either. What is journaled:: The records the device stores locally for later synchronization, and the worst-case outage the journal is sized to hold. What stops:: The functions that stop, and whether each one stops safely or waits to resume. What the operator sees:: The exact indication on the device, what the operator is told to do, and, for a device that forwards alarms, what the receiving system shows when the device goes silent.
The Offline Failure Mode Worksheet: one row for each dependency failure the device must survive.
| Dependency failure | Degraded mode | What continues | What is journaled | What stops | What the operator sees |
|---|---|---|---|---|---|
| Enterprise network down | |||||
| DNS fails | |||||
| Identity provider unavailable | |||||
| Vendor cloud unavailable | |||||
| Gateway unavailable | |||||
| Time service unavailable | |||||
| Credentials expired | |||||
| Software update incomplete | |||||
| Remote AI service unavailable |
Fill in every cell. An empty cell means nobody designed that degraded mode. A “What the operator sees” cell that says nothing is shown fails its row, because staff who don’t know the device is offline keep trusting records it isn’t sending.
Then record these answers for the device as a whole:
Journal capacity:: State the worst-case outage the journal holds. State the fill level at which the device alerts. List the low-value telemetry it sheds first. Confirm that it never sheds safety or therapeutic records. Describe what the operator sees if the device can no longer journal. Unsynchronized records:: Show the count and age of unsynchronized records on the device and in the fleet view. Describe the alert when that age crosses the limit the institution sets. Atomic write:: Show that each journal record and its pending-send state are written in one transaction, or as one appended record that carries its own state. Journal epoch:: Describe how the device starts a new epoch, with a new epoch or boot identifier, when its journal is wiped or replaced, so that no sequence number is reused. Clock quality:: Show that each record carries the device’s time, a time-quality flag, and a monotonic counter. Show that the device journals a clock-adjusted event whenever its clock is stepped. Describe how time sync measures the device-to-reference offset at reconnection. Command expiry:: Show that the device measures expiry by elapsed time on a monotonic clock since receipt, or by the receiver’s verified time, and never by an unverified device clock. Post-outage reconciliation:: Describe how records journaled during the outage arrive marked as delayed, go through the clinician review step, and are reconciled against what staff charted by hand. Confirm that no record overwrites a clinician’s entry silently.
Each answer is a design decision. Write each one down before the first release.
