Case: the RTO nobody had tested
The situation
A mid-sized insurer declares a 4-hour RTO for its claims management platform. The figure has sat in the BIA for five years, it is reproduced in the annual report to the regulator, and nobody has ever challenged it.
IT confirms: failover to the recovery site is automated and takes around 90 minutes under test conditions.
The exercise
A failover exercise is run on a Tuesday at 2 p.m., under conditions as realistic as the organisation can bear. Observed timeline:
| Time | Event | Elapsed |
|---|---|---|
| 14:00 | Simulated outage declared | 0 min |
| 14:18 | Monitoring confirms total unavailability | 18 min |
| 14:41 | On-call reached, initial diagnosis made | 41 min |
| 15:25 | Team assembled, failover decision taken | 85 min |
| 15:52 | Operations director's authorisation obtained | 112 min |
| 17:31 | Technical failover complete | 211 min |
| 18:47 | Functional tests validated, service reopened | 287 min |
Actual duration: 4 h 47. Target: 4 h. Overrun: 47 minutes.
What the exercise actually reveals
The technical failover took 99 minutes, in line with the estimate. The overrun does not come from the technology.
It comes from three items the RTO never counted:
- Detection and diagnosis: 41 minutes. The RTO clock starts at the incident, not at the decision.
- Decision: 71 minutes. Assemble the team, assess, arbitrate, obtain formal authorisation. On a Tuesday afternoon, with everybody available.
- Verification: 76 minutes. Nobody had anticipated that functional validation would take longer than the failover itself.
The decisions that followed
The insurer did not raise its declared RTO: it closed the gap.
- Standing decision delegation to second-level on-call below a defined impact threshold, removing the wait for authorisation;
- Automatic failover criteria: beyond 30 minutes of confirmed total unavailability, failover proceeds without a meeting;
- Reduced functional test set from 40 checks to 9 blocking checks, the other 31 being run after reopening;
- Two exercises a year, one of them outside business hours.
The next exercise, on a Saturday morning, measured 3 h 38. The 4-hour RTO became a commitment, where until then it had been a number.
Key takeaways
- An untested RTO is a hypothesis, not a commitment
- Decision time is almost always missing from estimates
- The measured gap is more useful than the target figure