Blog

A tabletop is not a DR test: what each one proves

Honest distinction between a BC/DR tabletop and a live failover or restore test: what each proves, where each helps, and why you still need both in a serious program.

Teams often blur "we ran a tabletop" with "we tested disaster recovery." Those are related readiness activities with different proofs. Confusing them produces weak evidence and overconfident status reports. This post draws a clean line so security, engineering, and auditors share the same vocabulary.

What a tabletop is

A tabletop is a facilitated discussion under a scenario. People in the room receive injects, make decisions in role, and produce a record of what they would do. For business continuity and regional outage themes, that often includes who declares an incident, who talks to customers, which systems are considered critical, and what order people think recovery should follow.

A tabletop is excellent at exposing unclear ownership, missing contacts, plan language that cannot be executed, and notification clocks nobody can meet. It is a people-and-process instrument.

What a DR test is

A disaster recovery test exercises technical recovery paths. Common forms include failover of a service to a secondary region, restore of a backup to a clean environment, or a controlled loss of a dependency with measured recovery time. The artifact is technical: did the system come back, how long did it take, what broke, what data was lost or replayed.

A DR test is excellent at exposing configuration drift, broken runbooks at the command level, RPO/RTO assumptions that fail in practice, and monitoring gaps during recovery. It is a systems instrument.

What each one proves

Tabletop evidence supports claims that leadership and operators practiced decision-making for a BC/DR scenario and that you recorded attendance, decisions, and plan gaps. It does not prove that failover scripts work or that last night's backup restores cleanly.

DR test evidence supports claims about technical recovery performance for the systems you actually exercised. It does not prove that executives know the customer notification path, or that legal and support were coordinated under time pressure, unless those people were part of a broader exercise design.

Where a BC/DR tabletop still helps

Even if engineering runs regular failover drills, a tabletop still helps when:

  • The written BC plan has not been walked with the people named in it
  • Customer and regulatory notification paths need practice under a clock
  • Cross-functional ownership (security, infra, product, support) is unclear
  • You need exercise evidence for a policy mandate that is about readiness process, not only restore metrics

Those outcomes are real. They are just a different kind of real than "secondary region took traffic."

Where you still need a real failover or restore test

If your risk is data loss, region loss, or dependency failure, you still need technical tests. No amount of discussion replaces restoring a backup, failing over a database, or verifying that secrets and DNS follow the runbook. If an auditor or customer asks for DR test evidence, hand them the technical record. Do not point only at a tabletop packet and hope the vocabulary slides.

How to talk about both without over-claiming

Say what you did in plain language: "We ran a facilitated BC/DR tabletop on date X; packet attached." Or: "We failed over service Y in staging on date Z; metrics attached." Do not say a tabletop "validated DR" or "proved recovery." Do not say a failover test "covers incident response training" unless the exercise actually included those roles and decisions.

Products and consultants that blur the line create liability for the buyer. Honest programs keep two lanes: process exercises with decision records, and technical tests with recovery metrics. ControlDrill runs the tabletop lane and produces evidence of that exercise. It does not replace a live failover or backup-restoration test, and it does not issue a compliance verdict.

Bottom line: use tabletops to practice and record human decisions under a BC/DR scenario; use DR tests to prove systems recover. Keep the labels accurate so your evidence stays useful when someone who was not in the room has to rely on it.