The RTO nobody had tested
A declared four-hour RTO, an exercise measured at four hours forty-seven, and three time items almost no plan counts.
The number and the reality
A mid-sized insurer has declared a four-hour RTO for its claims platform for five years. The figure sits in the BIA, is reproduced in the annual report to the regulator, and nobody has ever challenged it. IT confirms, moreover, that failover to the recovery site takes around ninety minutes under test conditions.
The first properly timed exercise measures four hours forty-seven.
Where the time goes
The technical failover took ninety-nine minutes, in line with the estimate. So the overrun does not come from the technology. It comes from three items the RTO never counted.
Detection and diagnosis: forty-one minutes. The RTO clock starts at the incident, not at the decision. Between the moment the service falls over and the moment somebody confirms it is a total outage, time elapses that technical estimates systematically ignore.
Decision: seventy-one minutes. Assemble the team, assess, arbitrate, obtain formal authorisation. On a Tuesday afternoon, with everybody available. On a Sunday at three in the morning, that item at least doubles.
Verification: seventy-six minutes. Nobody had anticipated that functional validation would take longer than the failover itself. Forty checks had been defined, all of them blocking.
What changed afterwards
The insurer did not raise its declared RTO. It closed the gap.
A standing decision delegation was given to second-level on-call below a defined impact threshold, removing the wait for authorisation. An automatic failover criterion was written in: beyond thirty minutes of confirmed total unavailability, failover proceeds without a meeting. The functional test set went from forty checks to nine blocking ones, the other thirty-one being run after reopening. And the rhythm moved to two exercises a year, one of them out of hours.
The next exercise, on a Saturday morning, measured three hours thirty-eight.
What the example generalises
Three lessons transpose almost everywhere.
First: an untested RTO is a hypothesis, not a commitment. Until somebody has timed it, the BIA figure describes an intention.
Second: decision time is almost always missing from estimates, because it belongs to no technical team. It is nonetheless the most compressible item, and it compresses through governance — written delegations, binding thresholds — not through investment.
Third: the measured gap is more useful than the target figure. A programme measuring four hours forty-seven then three hours thirty-eight demonstrates improved capability. Two reports concluding "conformant" demonstrate nothing.
Also worth reading
Why your BIA bogs down, and how to unblock it
A hundred and twenty questionnaires, six months, and a table where every activity is four-hour critical. The cause is a sequencing error.
The backup that survives a compromised administrator
"We're active-active" is not an answer to the ransomware question. Here is what is.
A successful exercise is one that reveals a gap
"Everything went well" is a negative signal. Here is how to design an exercise that produces usable information.