Disaster Recovery Failover Test Checklist Template
Most DR tests stop when the recovery site comes up and everyone relaxes. The clock was never started properly, nobody knows how much data was lost, and the trip back to production is improvised at midnight.
This free disaster recovery test checklist is the runbook for one planned test, used by infrastructure leads, DR coordinators and MSPs testing on a client’s behalf. It takes the test from objectives and a go/no-go decision through failover, validation and failback, then measures real recovery time and data loss against each system’s targets. It covers a live failover and an isolated drill that leaves production running. The output is a timestamped test record, an issues log, updated runbooks and a signed report.
A Runbook for One Test, Not an Assessment of the Programme
NIST SP 800-34 Rev. 1, the US government’s contingency planning guide, separates tests, which use measurable results to prove a system works, from exercises, which simulate an emergency. For high-impact systems it expects a full-scale functional exercise that includes failover to the alternate location and a full recovery and reconstitution to a known state. That is the event this checklist runs.
AWS Elastic Disaster Recovery draws a useful line too. It calls a non-disruptive test a drill, and treats failover, redirecting production traffic to the recovery instances, as a step you perform outside the service. Recovering machines and moving users are separate jobs, and the test mode chosen on the first task decides which of them this run covers.
Backup restore test
Can this backup be restored?
Scope: a rotating sample of systems restored into an isolated network.
If nothing is switched over and the team talks the scenario through, it is a tabletop, not a failover test. Business tabletops and functional exercises run through the Business Continuity Plan Testing Checklist, which leaves the IT recovery test to IT operations. This page is that test.
What the DR Failover Test Checklist Covers
Seven phases take one test from objectives to a signed report. The test mode decides which failover and failback tasks appear, the test authority approves the go/no-go before anything moves, and the DR plan owner signs off the result. There is no fixed schedule: open one record per test from your DR test programme.
Phase 1
Phase 1: Scope, Objectives & Roles
The test mode recorded on the first task, Isolated drill or Live failover, decides which tasks appear in Phases 2, 4, 5 and 6.
Open the test record and name the test lead, test authority and scribe — with the test mode, the date, who can stop the test and the DR plan owner who signs off
List the systems in scope in recovery order — each with its RTO and RPO taken from the DR plan, and the application owner who will check it
Write the objectives and success criteria — for example, orders can be taken at the recovery site within four hours with no more than 15 minutes of data lost
Choose the scenario and state what is out of scope — loss of the primary site, a cloud region or a storage array, and the systems nobody may touch
Record the runbook version the team will follow — the test checks the document as well as the technology, so nobody works from memory
Phase 2
Phase 2: Readiness & Dependencies
The DNS task appears only for a live failover.
Confirm replication is healthy and inside the RPO for every system — replication lag, the latest consistent recovery point and any paused or failing jobs
Check the recovery site has the capacity, licences and configuration it needs — quotas, licence keys, patches and firewall rules drift between tests
Confirm identity, DNS and certificates work at the recovery site — directory services, name resolution and expiring certificates stop more recoveries than servers do
Shorten DNS TTLs on the records that will move — a day ahead, so users follow the switch in minutes and follow it back at failback
Raise the change record and freeze unrelated changes — a patch or deployment in the same window makes the results impossible to read
Brief users, the service desk, on-call staff and suppliers — the window, the expected impact and who to call if something unexpected breaks
Phase 3
Phase 3: Go/No-Go
The test authority named in Phase 1 records Approved or Not approved on the last task. The checklist halts there, and nothing fails over until they decide.
Agree the abort criteria and the way back — for example, stop if recovery is not complete at one and a half times the RTO, and how production is restored if it is
Record a reference point in production before the test — row counts on key tables, the last transaction number and the time it was written
Confirm every runbook step has an owner on shift with working access — including break-glass accounts and the passwords stored outside the primary site
Check for open incidents and new risks on the day — a live major incident or a degraded storage array is a reason to postpone
Approve or postpone the test — the test authority records the decision and the reason, with the readiness checks attached
Phase 4
Phase 4: Failover
Stopping writes at the primary and redirecting users appear only for a live failover. In a drill, production keeps running.
Start the test clock and log the declared start time — every later time is measured from this one
Stop writes at the primary and record the final replication point — so no transaction lands in both sites and the data loss can be measured
Bring up the recovery copies in the documented order — identity and databases first, then application tiers, then front ends, logging each start time
Redirect users and integrations to the recovery site — DNS, load balancers, remote access and the partner endpoints that call in
Log every departure from the runbook as it happens — the step, the time, what was done instead and who decided
Record the time each system is declared recovered — when its application owner can use it, not when the virtual machine booted
Phase 5
Phase 5: Validation at the Recovery Site
Application owners picked in Phase 1 run their own checks. The isolation check appears only for a drill.
Run the functional checks for each system — sign in, open recent records, complete a transaction and run a report
Measure how current the recovered data is — compare the newest record at the recovery site with the Phase 3 reference point
Test integrations, scheduled jobs and outbound email — the parts nobody notices are missing until month-end
Confirm monitoring, backup and security tools cover the recovery site — a system running unwatched and unprotected is not recovered
Prove the drill environment cannot reach production — no duplicate hostnames on the live network, and no job sending real emails, orders or payments
Record each application owner’s acceptance — Accepted, or Not accepted with the open issues listed
Phase 6
Phase 6: Failback & Return to Normal
The first three tasks appear only for a live failover. Every test ends with replication healthy and the temporary resources gone.
Agree the failback window and announce it — failback is a second change, with its own downtime and its own go decision
Reverse replication from the recovery site to the primary — every change written during the test has to come home before users do
Switch users back and confirm the primary holds every change — compare against the final state of the recovery site before reopening
Confirm replication to the recovery site is healthy again — and take a fresh full backup; DR protection is reduced until both are done
Remove drill instances, test accounts and temporary access — they hold live data, cost money and widen the attack surface
Log the time normal operations resumed and tell users — this stops the test clock
Phase 7
Phase 7: Results, Report & Sign-off
Answering No to “All targets met?” adds the retest task. The DR plan owner signs off the report, and the checklist halts until they record a decision.
Calculate the measured recovery time and data loss for each system — from the logged timestamps, against its RTO and RPO
Log every issue with an owner and a due date — from the deviation log, the validation checks and the hot debrief held straight after
Schedule a retest for each missed target — after the fix, and before anyone quotes the RTO to the business again
Update the runbooks and the DR plan — the order, steps, contacts and timings that turned out to be different
Write the test report — objectives, timeline, results against targets, issues and lessons learned: the after-action report NIST SP 800-34 asks for
Sign off the test report — the DR plan owner records Approved or Not approved, with the reason
A recovery time worked out afterwards from memory is always flattering. The scribe logs these points as they happen, and the results in Phase 7 are simple subtraction.
Timestamp
Logged when
Used for
Declared start
The test lead starts the clock in Phase 4
The zero point for every recovery time
Last write at the primary
Writes stop, or the reference point in a drill
The start of the data loss window
Newest recovered record
Checked at the recovery site in Phase 5
Data loss: last write minus newest recovered record, against RPO
System declared recovered
Its application owner can use it
Recovery time: this minus declared start, against RTO
Users redirected
DNS and load balancers point at the recovery site
How long users waited after the system was ready
Normal operations resumed
Failback complete, users back on the primary
The total disruption the test caused
Drill or live failover
An isolated drill brings the recovery copies up on a separate network while production carries on. Microsoft describes an Azure Site Recovery test failover as a drill that does not affect ongoing replication or production, and recommends an isolated test network. If a drill runs in the production network instead, shut the primary down first: two machines with the same identity on one network cause unexpected results.
A drill measures how quickly systems come up and how current their data is. It cannot measure user redirection, real load at the recovery site or failback. A common pattern is drills through the year and one live failover and failback per critical system each year, but NIST leaves the frequency organisation-defined.
Why Run DR Tests in CheckFlow?
1
A timeline that writes itself
The activity trail records who completed each step and when, so the failover timeline comes from the record, not from the scribe’s notes. A table inside the systems task holds each system’s targets and measured results, and comments capture decisions made under pressure.
2
Nothing moves without a decision
The go/no-go goes to the test authority picked on the first task, and the checklist halts until they record Approved or Not approved. Conditional logic shows the live failover and failback tasks only when the test mode calls for them, and the drill isolation check only for a drill.
3
Findings that reach the next test
Each issue is assigned to a person or group with a due date. Validation screenshots and the report attach as audit evidence, and template versioning keeps a history as the runbook improves after each test.
Our disaster recovery checklist guide covers the programme around this test: setting RTO and RPO, the business impact analysis and the range of test types. The annual Disaster Recovery Audit Checklist uses this checklist’s signed reports as its evidence that failover was tested.
It is a planned event in which systems are recovered at the secondary site or cloud region, used there for a period, and then returned to the primary. Unlike a single restore test, it exercises the whole recovery plan, and it measures real recovery time and data loss against the targets the business was promised.
What is the difference between failover and failback?
+
Failover moves the service from the primary site to the recovery site. Failback brings it home again, which means replicating every change made at the recovery site back to the primary before users switch over. It is rehearsed less than failover, so treat it as a second change with its own window and go decision.
How do you measure RTO and RPO during a DR test?
+
Log timestamps as the test runs. Measured recovery time is the time a system was declared usable by its owner minus the declared start of the test, compared with its RTO. Data loss is the time of the last write at the primary minus the time of the newest record found at the recovery site, compared with its RPO. Note the last transaction number beforehand so the comparison is exact.
Can you test disaster recovery without downtime?
+
Yes, with an isolated drill. Tools such as Azure Site Recovery and AWS Elastic Disaster Recovery can bring up recovery copies while replication and production carry on, ideally on a network that cannot reach production. A drill does not prove user redirection or failback, so the most critical systems still need a live failover from time to time.
Is NIST SP 800-34 still current?
+
Yes. Revision 1, last updated in November 2010, is still the version NIST lists as final in October 2026, and no Revision 2 has been published. It ties the depth of testing to a system’s impact level: a tabletop for low-impact systems, a functional exercise with recovery from backup for moderate, and a full-scale functional exercise with failover to the alternate site for high-impact systems.
Is CheckFlow free for this template?
+
14-day free trial, no card required. The Business plan is $10 per user per month after the trial. Full details at checkflow.io/pricing.
Prove the Recovery Time Before You Need It
Free trial — no credit card required.
Do you like cookies? 🍪 We use cookies to ensure you get the best experience on our website. Learn more