Back to Blog
Business Continuity
9 min read

Disaster Recovery for Hospitals: Designing for EHR Outages

RTO and RPO are commitments, not aspirations. Setting them per clinical service, designing replication that ransomware cannot follow, and sequencing a restore that actually works.

GuardsArm Team

Security Experts

November 30, 2025

Disaster recovery design

Disaster recovery is the engineering problem: what infrastructure exists so that a failed clinical system can be brought back, how quickly, and losing how much data. It is distinct from continuity — covered in the 72-hour plan — and conflating the two is why hospitals end up with a replication design and no ability to treat patients during the outage.

Per service
A single hospital-wide RTO is meaningless — the ED and the payroll system differ enormously
Replication is not backup
Synchronous replication copies corruption and encryption faithfully
Sequence matters
Restoring in the wrong order produces systems that start and then fail

Set objectives per clinical service

Asking "what is our RTO" produces a number nobody can honour. Ask it per service, with the clinical owner in the room:

ServiceRealistic RTORPOWhy
EHR read access1-4 hoursNear zeroSafe care needs allergies and medications
EHR write / ordering4-12 hoursMinutesPaper can bridge, briefly
PACS / imaging4-12 hoursNear zeroDiagnostic dependency; images cannot be recreated
Laboratory system2-8 hoursMinutesResults drive immediate decisions
Pharmacy2-8 hoursMinutesMedication safety
Scheduling24 hoursHoursDisruptive, not immediately unsafe
Payroll / financeDaysDayGenuinely tolerable

The exercise of writing this table is most of the value. It forces the conversation about what the organisation will actually pay to protect, rather than asserting that everything is critical.


Replication copies your problems faithfully

Synchronous replication protects against hardware and site failure. It offers no protection at all against ransomware, corruption or a bad change, because it replicates those too — instantly, by design.

Layers of recovery capabilityEach layer protects against a different failure mode; only the immutable copy protects against an attacker with administrative access.Immutable / air-gapped copyCannot be altered or deleted within the retention window — this is the ransomware controlOffsite backupSeparate location, separate credentials, tested restoresAsynchronous replicaLag provides a small window to stop corruption propagatingSynchronous replicaSite and hardware failure only — copies encryption immediatelyProductionThe thing being protected
An estate with only the bottom two layers has high availability and no ransomware recovery.

The distinction that matters: can an attacker holding domain administrator credentials destroy this copy? If yes, it is availability infrastructure, not recovery infrastructure. Immutability and separate credentials are what move a copy into the top tier.


Restore sequencing

Clinical systems have dependency chains, and the order is not obvious under pressure. Write it down before you need it.

Restore sequence for clinical systemsNetwork and DNS first, then identity, then the data tier, then applications, and finally the interfaces that make the applications a clinical service.Network + DNSnothing works without itIdentitydirectory, then authData tierdatabases, consistency-checkedApplicationsEHR, PACS, LISInterfacesHL7/FHIR — the service is not back without these
Interfaces last and most forgotten — a restored EHR that cannot receive lab results is not restored.
Identity is the dependency people forget
If the directory is part of the outage, nothing else authenticates. Domain controller recovery needs its own tested procedure, its own credentials stored outside the environment, and a documented break-glass account that does not depend on the directory being up.

What to verify annually

  • Recovery of a domain controller, from backup, in isolation
  • Database restore with a consistency check, not just a successful copy
  • That the immutable copy genuinely cannot be deleted by an administrator
  • That recovery credentials are available when the primary credential store is down
  • That the documented sequence matches reality — run it and time it
  • That interfaces come back and exchange real messages

Where to start

Pick your most critical clinical system and answer one question: if an attacker had domain administrator rights for an hour, which copy of its data would survive? If the answer is "none", that is the gap, and it is a funding conversation rather than a process one.

For whether those copies restore, see backup testing.

GuardsArm assesses clinical recovery architecture including immutability verification and restore sequencing. Book a scoping call.

Written by GuardsArm Team

Our team of cybersecurity experts brings decades of combined experience in penetration testing, compliance auditing, and incident response. We're dedicated to helping organizations strengthen their security posture.

Take the next step on “Disaster Recovery for Hospitals: Designing for EHR Outages”

Talk to the GuardsArm team about how these services apply to your environment.