What it protects against
Power failure, water, fire and earthquake, all rare and all possible. And ransomware, where a properly thought-through disaster recovery plan is the difference between days and weeks.
Where we start
With two numbers. RPO, how much data loss is acceptable, and RTO, how long it may take to be running again. Both have to be verified and signed off with the stakeholders, because they carry the whole architecture and are not a technical quantity.
After that we classify and prioritise data and systems. Not everything needs the same number, and trying to treat everything alike is the most common reason a plan becomes unaffordable.
How the design comes together
It accounts for the workloads in your datacentres and in your public clouds, along with availability and performance requirements. Site choice feeds in too: where your datacentres and cloud regions sit partly determines what is possible at all.
Implementation, and the part that gets skipped
We build on Zerto and Veeam. Both allow the real event to be rehearsed: virtual servers can be started in a sandbox at the recovery site and checked without touching production.
That is exactly the part that gets skipped. Regular test scenarios with the system administrators and the affected departments are what separate a plan from a hope.