Before you start
Backup console access, somewhere isolated to restore into, and a booked slot with the service owner.
Agree the scope and recovery plan with the service owner. Adapt the exercise to your environment.
Choose a service and define success
Agree the scope with its owner and the people responsible for the recovery platform. Write the ordinary business task you will use to judge recovery, the expected recovery time and the acceptable age of recovered data. If these expectations have never been agreed, record that gap before interpreting the test. Choose a representative service whose failure you can safely simulate; a first drill need not include every dependency.
List the dependencies before reserving the window
Identify the backup, encryption-key access, software versions, configuration, certificates and identity services needed to use the recovered data. Record the approved route to each, without copying secrets into the test plan. Confirm the required people will be available and agree a stop condition: for example, an unexpected connection to production or uncertainty about the destination. Reserve additional time for investigation and cleanup.
Prove the target is isolated
Ask the responsible network operator to confirm the restore target cannot contact production services or send live messages, payments or scheduled jobs. Check approved connectivity tests from the target before restoring sensitive data. Disable or redirect external integrations through a documented method appropriate to the application. If you cannot demonstrate isolation, stop and resolve that first. An environment named “test” is not evidence that its routes are safe.
Restore the agreed recovery point and keep a timeline
Record the backup identifier and recovery-point timestamp. Start the clock at the first recovery action, including retrieving access and preparing the target. Follow the service runbook and log each manual decision, error and missing dependency. Do not silently substitute a newer backup when the chosen one fails: record the failure and agree whether a second attempt is in scope. Use vendor-specific recovery instructions for the actual product.
Test a business task and inspect the data
Have the owner perform the agreed task against the isolated copy, using approved test inputs. Check relevant record counts, recent transactions and relationships as appropriate to the service. Record when the task succeeds and how recent the recovered information is. A fictional booking system might load successfully but be missing the latest reservations; that difference belongs in the result even if the restore software reports success.
Turn the findings into assigned work
Compare the result with the agreed recovery expectations. Separate measured observations from estimates: if an unavailable dependency was simulated, the full service recovery time remains unproven. Assign each gap an owner and a due date. For a failed requirement, record whether the next step is technical work, a revised recovery design or a business decision about the required capability. Retest material fixes.
Remove the test copy and retain useful evidence
Dispose of restored data, temporary accounts and resources using the agreed process, and confirm cleanup with the responsible operator. Keep the backup identifier, timeline, isolation evidence, acceptance checks and action list in the approved service record. Book the next drill with a named owner. Scale the test frequency and scope to the service and its changes rather than treating one successful exercise as permanent assurance.
What you should leave with
A dated recovery report with the tested recovery point, elapsed time to a usable service, data checks, untested dependencies and assigned follow-up actions. The owner should be able to explain what the test established and what it did not.