Latest News

Data Centre Power Outage Recovery: A Practical First-Hour Runbook

Use this practical first-hour runbook to coordinate safety, UPS, generators, IT recovery, communications and evidence after a data centre outage.
26 August 2026
Data Centre Power Outage Recovery: A Practical First-Hour Runbook

A data centre power outage is rarely solved by one team. Facilities must understand the electrical state, IT must protect services and data, security must control access, and leadership needs reliable information. If those workstreams are not coordinated, well-intended actions can compete with each other.

This first-hour runbook is a planning template, not a switching instruction. Every site needs its own approved procedures, single-line diagrams, alarm matrix and competent authorised people. During an event, safety and the site's rules take priority over the clock.

The first principle: stabilise before you restore

The goal is not to re-energise everything as quickly as possible. It is to establish a safe known state, protect the remaining power path, preserve evidence and restore services in a controlled order. An unstable partial recovery can cause a second interruption and complicate diagnosis.

Before minute zero: what the runbook should already contain

  1. named incident roles and deputies for facilities, IT, communications and leadership
  2. approved electrical operating procedures and escalation boundaries
  3. current single-line diagrams and critical-load schedules
  4. UPS, generator, ATS and fuel supplier contact details
  5. service priorities, recovery dependencies and shutdown sequences
  6. alternative communications if normal systems are unavailable
  7. decision logs, event-record forms and secure evidence storage
  8. pre-agreed criteria for temporary power, controlled shutdown or site evacuation

Minutes 0-5: protect people and declare the incident

Confirm immediate safety. Check for fire, smoke, smell, heat, flooding, electrical damage or other hazards without entering restricted areas. Contact emergency services where required and follow site evacuation rules.

Declare the incident through the agreed route and appoint one incident lead. Record the time, who reported it and the first observable facts. Avoid conclusions such as 'generator failure' until the evidence supports them.

  • Is the loss total, partial or limited to one power path?
  • Are any people or areas at immediate risk?
  • Which safety and security systems remain available?
  • Who is the incident lead and where is the decision log?
  • Which technical responders are being called?

Minutes 5-15: establish the power state

Facilities should build a concise status picture from displays, alarms and remote monitoring. Record, rather than repeatedly reset, the UPS state, battery runtime estimate, bypass availability, generator state, ATS position, fuel information and environmental alarms.

  1. utility supply present, absent or unstable
  2. UPS normal, on battery, bypassed, overloaded or shut down
  3. battery runtime shown and trend over the last few minutes
  4. generator available, starting, running, failed or carrying load
  5. ATS or switchgear position and any inhibited transfer
  6. room temperature, cooling status and water-leak alarms
  7. current critical load and any unexpected demand

Photograph displays and preserve event logs where policy allows. Use exact alarm text and timestamps. Do not open live equipment or operate switchgear outside the approved authority and procedure.

Minutes 5-15: establish the service state

In parallel, IT should identify which services are healthy, degraded or unavailable. Check monitoring from more than one point because an apparently healthy server does not prove that network, storage, authentication or external connectivity is working.

Link service status to power dependencies. If one UPS path has failed, determine which racks, network zones or storage systems it supplies. This is where an accurate critical-load map saves time.

Minutes 15-30: choose the operating strategy

With a known power and service state, the incident lead can choose among prepared strategies: hold on the stable power path, reduce non-essential load, begin controlled shutdown, restore a proven supply, or prepare temporary generation. The decision should consider remaining autonomy, fuel, cooling, staff access, fault uncertainty and the time needed for each action.

Set explicit decision points. For example, if battery runtime falls to the site's approved trigger before generator support is proven, the controlled-shutdown plan begins. The exact thresholds must be set by the site in advance.

Control load before adding load

Do not reconnect all services at once. Large load steps, transformer energisation, motor starts and battery recharge can place additional demand on generators and distribution. Use the site's staged restoration sequence and verify stability after each agreed step.

  • retain life-safety, control and core network services
  • remove or defer non-essential test, development and office loads
  • delay high-inrush mechanical loads until the source is stable
  • coordinate UPS battery recharge with generator capacity
  • monitor voltage, frequency, load and temperature as services return

Minutes 30-45: recover services by dependency

IT recovery should follow the dependency map, not the loudest request. Core network, identity, storage, virtualisation and databases may need to recover before applications. Validate data integrity and application state rather than treating power-on as service restoration.

Assign an owner to each recovery group and record start, result and exceptions. If a step fails, stop and reassess rather than repeatedly cycling equipment.

Minutes 45-60: confirm stability and communicate

By the end of the first hour, leadership needs a short factual update: safety status, power source, remaining constraints, customer or operational impact, actions completed, next decision point and help required. State uncertainty openly.

  • current source and whether redundancy is degraded
  • critical services restored, degraded or offline
  • battery runtime, fuel autonomy and cooling constraints
  • technical attendance and estimated next assessment time
  • next controlled action and its risk
  • time of the next stakeholder update

Evidence to preserve

Keep UPS, generator, ATS, BMS and monitoring logs aligned to a common timeline. Retain display photographs, service alerts, switching records, load trends, fuel data, IT monitoring and decision logs. Evidence supports root-cause analysis and prevents memories from becoming the only record.

What not to do in the first hour

  • do not improvise switching or defeat interlocks
  • do not reset every alarm before it is recorded
  • do not assume a running generator is carrying the required load
  • do not restart services without checking dependencies and source capacity
  • do not let several teams issue conflicting status messages
  • do not postpone a controlled shutdown until no safe margin remains

After stabilisation: recover resilience, not only service

A site can be operational while still running without redundancy. Record every degraded component, temporary bypass and deferred load. Agree what must be repaired or tested before the incident is closed and what contingency remains in place meanwhile.

Complete a post-incident review while evidence is fresh. Identify technical cause, response delays, unclear authority, missing spares, monitoring gaps and documentation errors. Update and exercise the runbook so the next response is stronger.

Frequently asked questions

Protect people, declare the incident, appoint an incident lead and establish the actual power and service state from recorded evidence.

Not automatically. Confirm that the source is stable and follow a staged, dependency-led restoration plan that considers load steps and data integrity.

At minimum, facilities or electrical operations, IT service owners and an incident lead. Security, communications, vendors and leadership may also be required.

Exact alarms and times, operating mode, input/output/bypass state, load, battery charge and runtime estimate, temperature and recent events.

At the pre-agreed decision point that preserves enough margin to complete it safely. The trigger depends on load, runtime, generator status and the site's recovery design.

Yes. It may be unloaded, overloaded, unstable, low on fuel or supporting only part of the distribution. UPS and cooling constraints can also remain.

The sequence and timing help distinguish cause from consequence and support safe diagnosis, warranty discussions and post-incident improvement.

Use a risk-based exercise programme and test after material changes. Include technical, decision-making and communications elements, not only document review.

Turn the first hour into a rehearsed sequence

A useful runbook connects power state, service dependencies, decision authority and communications. It should be site-specific, accessible during an outage and tested before it is needed.

Contact P&I Group to review data-centre standby generation, UPS resilience, maintenance or emergency power contingency.