Case study12 min

Case: the RTO nobody had tested

The situation

A mid-sized insurer declares a 4-hour RTO for its claims management platform. The figure has sat in the BIA for five years, it is reproduced in the annual report to the regulator, and nobody has ever challenged it.

IT confirms: failover to the recovery site is automated and takes around 90 minutes under test conditions.

The exercise

A failover exercise is run on a Tuesday at 2 p.m., under conditions as realistic as the organisation can bear. Observed timeline:

TimeEventElapsed
14:00Simulated outage declared0 min
14:18Monitoring confirms total unavailability18 min
14:41On-call reached, initial diagnosis made41 min
15:25Team assembled, failover decision taken85 min
15:52Operations director's authorisation obtained112 min
17:31Technical failover complete211 min
18:47Functional tests validated, service reopened287 min

Actual duration: 4 h 47. Target: 4 h. Overrun: 47 minutes.

What the exercise actually reveals

The technical failover took 99 minutes, in line with the estimate. The overrun does not come from the technology.

It comes from three items the RTO never counted:

  1. Detection and diagnosis: 41 minutes. The RTO clock starts at the incident, not at the decision.
  2. Decision: 71 minutes. Assemble the team, assess, arbitrate, obtain formal authorisation. On a Tuesday afternoon, with everybody available.
  3. Verification: 76 minutes. Nobody had anticipated that functional validation would take longer than the failover itself.

The decisions that followed

The insurer did not raise its declared RTO: it closed the gap.

  • Standing decision delegation to second-level on-call below a defined impact threshold, removing the wait for authorisation;
  • Automatic failover criteria: beyond 30 minutes of confirmed total unavailability, failover proceeds without a meeting;
  • Reduced functional test set from 40 checks to 9 blocking checks, the other 31 being run after reopening;
  • Two exercises a year, one of them outside business hours.

The next exercise, on a Saturday morning, measured 3 h 38. The 4-hour RTO became a commitment, where until then it had been a number.

Key takeaways

  • An untested RTO is a hypothesis, not a commitment
  • Decision time is almost always missing from estimates
  • The measured gap is more useful than the target figure