30. The City That Planned for Ruin

Domain: Business Continuity, Disaster Recovery — BIA, RTO, RPO, Backups POV: Kip Reading time: ~15 minutes


The eastern gear-chamber breach had demonstrated that the city could survive a disaster. The question that Kip could not stop asking — the question that had been growing in the back of his mind since the first gate failure, since the first whisper from the Wall — was whether the city could survive the worst disaster imaginable. And if it could, how long it would take to recover, and what would be lost in the meantime.

“We need a continuity plan,” Kip said to the council. “Not just incident response — that is about the first six hours. We need a plan for the six days, the six weeks, the six months after a catastrophic failure. We need to know what the city’s most critical functions are, how quickly they must be restored, and how much data we can afford to lose.”


The business impact analysis — a term that Quill had found in the old treatises and that Kip had adopted without fully understanding its etymology — was the most comprehensive assessment the Citadel had ever attempted.

Every critical function of the city was identified: the Great Clock that regulated time, the gear-trains that powered the infrastructure, the water systems that supplied the districts, the food distribution network that kept the citizens alive, the communication channels that connected the Citadel to the Archipelago and the Weave-lands. For each function, the analysis asked three questions.

First: what is the recovery time objective — the maximum amount of time that the function can be unavailable before the consequences become unacceptable? The Great Clock, for example, could tolerate a disruption of perhaps four hours before the desynchronization of the gear-trains caused cascading failures across the city. The food distribution network could tolerate longer — perhaps two days, if emergency stockpiles were available. The communication channels to the Archipelago could tolerate a full week before the diplomatic consequences became severe.

Second: what is the recovery point objective — the maximum amount of data that can be lost without causing irreparable harm? The Archive of Record’s logs, for example, could tolerate the loss of perhaps an hour of data — the events during that hour could be reconstructed from secondary sources. The Library’s catalog of classified documents could not tolerate any loss at all — every document’s classification metadata had to be preserved perfectly, or the entire classification system would be compromised.

Third: what is the backup strategy — how will the data and the systems be preserved so that they can be restored after a disaster? This was the question that Kip found most difficult, because the old city had almost no backup strategy at all. The Archive of Record kept a single copy of everything. The Library’s catalog was maintained in a single ledger. The gear-chamber configurations were documented in the engineers’ notebooks, which were themselves stored in the gear-chambers — a single point of failure that a determined adversary could destroy with a single attack.

“We need backups,” Kip said, and the word felt simultaneously obvious and revolutionary. “Multiple copies of everything critical, stored in different locations, on different media, updated at different intervals. A full backup of the Archive’s logs every week, stored in a vault beneath the Spire. An incremental backup of the Library’s catalog every day, stored in a separate facility in the Archipelago. A real-time replication of the gear-chamber configurations, updated continuously, so that if one set of configurations is destroyed, another is immediately available.”

“Immutable backups,” Thread added quietly. “Copies that cannot be altered — even by someone with administrator access. The adversary has shown that they can corrupt any system they can access. The backups must be protected not just from physical destruction but from tampering. Once written, they cannot be changed.”


The backup strategy was tested — inevitably, and sooner than anyone expected — when a fire broke out in the Archive of Record.

The fire was not an attack — at least, the investigation suggested it was not. An old heating mechanism in one of the Archive’s storage rooms had malfunctioned, and the resulting fire had destroyed several shelves of records before the Archive’s suppression systems extinguished it. The records were among the oldest in the Archive — tax rolls and birth ledgers and treaty scrolls dating back to the Citadel’s founding. Irreplaceable, except for the fact that they had been backed up.

The backups — stored in the vault beneath the Spire, on media that had been designed to survive fire and flood and the corrosive effects of the Shroud itself — were retrieved and restored. The process took three days, because the backup restoration procedures had not been practiced and the archivists had to improvise parts of the process. But it worked. The records were recovered. The Archive was restored to its pre-fire state, with a data loss of less than six hours — the interval between the last backup cycle and the moment the fire suppression systems activated.

“The recovery worked,” Sable said, reviewing the restoration report. “But it took three days. The recovery time objective for the Archive’s critical records was supposed to be one day. We missed the target by two days.”

“Because we had not practiced,” Kip said. “The backup strategy is only half the plan. The other half is the restoration procedure — the steps required to retrieve the backups, verify their integrity, and restore them to the production system. And the restoration procedure must be practiced. Regularly. Under realistic conditions. Because the first time you attempt a restoration cannot be during an actual disaster.”

“The treatises call it disaster recovery testing,” Quill said. “Scheduled exercises that simulate a catastrophic failure and require the response teams to restore critical functions within the recovery time objectives. The exercises are not just about verifying that the backups work — they are about training the people who will need to execute the restoration, and identifying the gaps in the procedures, and building the muscle memory that will be needed when the real disaster arrives.”


The first full-scale disaster recovery exercise was conducted one month after the Archive fire. The scenario was the worst that Kip could imagine: a coordinated attack that destroyed the Great Clock, severed the primary gear-trains, and corrupted the Spire’s timekeeping infrastructure. The city’s recovery time objectives required that essential timekeeping be restored within four hours and full synchronization within twenty-four.

The exercise revealed gaps that no amount of planning had anticipated. The backup clock mechanisms — secondary timekeeping devices that had been installed in each district — were functional but had not been synchronized recently, and their readings diverged by as much as two minutes. The restoration procedure for the Spire’s master clock assumed that the physical clock-face would be intact, but the exercise scenario involved physical destruction, and the procedure had no provision for rebuilding the clock-face from scratch. The communication plan for notifying the citizens of the timekeeping disruption was incomplete — it assumed that the city’s public announcement system would be functional, but the scenario involved the destruction of the Spire, which housed the announcement system’s primary relay.

The exercise was, in short, a partial failure. But it was a useful failure — far more useful than a successful exercise would have been, because it identified the gaps and forced the fellowship to address them. The backup clocks were re-synchronized and a procedure was established for regular synchronization checks. The Spire restoration procedure was expanded to include provisions for physical reconstruction. The communication plan was revised to include multiple redundant channels, so that the failure of any single channel would not prevent the citizens from receiving critical information.

“The exercise failed,” Kip said to the council, in the debrief that followed. “But the concept succeeded. We learned what we needed to learn. The next exercise will be better. And the one after that will be better still. Disaster recovery is not about getting it right the first time. It is about practicing until you cannot get it wrong.”

He looked at the backup vault beneath the Spire — the shelves of stored data, the replicated configurations, the immutable copies that could not be altered or destroyed — and he understood something that he had not understood before. The backups were not just data. They were hope — a bet that whatever was lost could be restored, whatever was destroyed could be rebuilt, whatever the adversary took could be taken back. The city had been fighting for its survival for months, and the fight would continue for months more. But the city had a plan now — not just for defending against attacks, but for recovering from them. And that plan, more than any wall or barrier or cryptographic system, was what would keep the city alive through the worst that was still to come.