DOWNTIME IS MORE EXPENSIVE THAN YOU THINK
The direct revenue loss from a data centre outage is significant — but it is the starting point, not the ceiling. Add SLA penalties, emergency engineering costs, hardware damage caused by abrupt shutdown, and the cost of rebuilding trust with affected customers, and the true figure can be multiples of the headline number.
For colocation providers, a single high-profile outage can trigger customer churn that takes years to recover from. For enterprise data centres, the knock-on impact to trading systems, payment platforms, and customer-facing applications compounds rapidly with every passing minute. Power failure is the single largest cause of unplanned data centre downtime — and it is the most preventable.
THE DATA CENTRE POWER CHAIN
Reliable data centre power is not a single product — it is a layered chain of infrastructure, each component protecting against a different failure mode:
- Utility grid input — the primary power source, subject to outages, voltage fluctuations, and quality issues outside your control.
- Automatic Transfer Switch (ATS) — detects grid failure and initiates switchover to backup generation within milliseconds.
- Uninterruptible Power Supply (UPS) — bridges the gap between grid failure and generator synchronisation, typically 10–30 seconds, while conditioning power quality.
- Backup generators — industrial gensets or gas turbines that sustain full facility load for hours or days during extended outages.
- Power Distribution Units (PDUs) — deliver conditioned, metered power to individual server racks with granular monitoring and control.
- Remote monitoring platform — provides real-time visibility of every component, enabling predictive maintenance and instant fault response.
Each link in this chain must be correctly specified, integrated, tested, and maintained. A weakness at any point can propagate through the entire system.
DATA CENTRE TIERS AND POWER REQUIREMENTS
The Uptime Institute's Tier Classification System defines four levels of data centre reliability, each with specific power infrastructure requirements:
Tier I — Basic
Single path for power and cooling. No redundancy. Minimum 99.671% uptime. Suitable for non-critical workloads only.
Tier II — Redundant
Redundant components on a single path. 99.741% uptime. Basic generator backup. Suitable for departmental systems.
Tier III — Concurrently Maintainable
Multiple active power paths. N+1 redundancy. 99.982% uptime (<1.6 hrs/yr). Industry standard for commercial colocation.
Tier IV — Fault Tolerant
Multiple active power paths simultaneously. 2N redundancy. 99.995% uptime (<26 min/yr). Required for hyperscale and financial infrastructure.
The gap between Tier II and Tier III is not incremental — it is the difference between a facility that can tolerate a power system failure and one that cannot. Specifying the right tier from the outset prevents the exponentially more expensive retrofit later.
THE GROWING CHALLENGE OF POWER DENSITY
Traditional data centres were designed for 5–10 kW per rack. Modern high-performance compute deployments — driven by AI training, GPU clusters, and high-frequency trading infrastructure — regularly demand 30–60 kW per rack, with next-generation installations pushing beyond 100 kW.
This has profound implications for power infrastructure. Generators sized for legacy load profiles are inadequate. UPS systems designed for conventional IT equipment cannot handle the step-load characteristics of GPU clusters. Cooling and power distribution systems become the binding constraint on density.
"Facilities that were adequately powered at build are increasingly discovering they've hit a hard ceiling. Retrofitting power infrastructure at hyperscale is not just expensive — it requires taking systems offline that cannot afford to go offline."
Planning for density growth from day one — with modular, scalable generator and UPS configurations — is the only way to avoid this constraint.
WHAT BEST-IN-CLASS LOOKS LIKE
The highest-performing data centre operators share a common approach to power infrastructure: they engineer for the failure, not just the operation. This means N+1 or 2N redundancy at every critical system, monthly and annual load testing under simulated fault conditions, continuous remote monitoring with automated alerting, and a maintenance programme that treats every genset and UPS module as mission-critical — because it is.
The Power Vault Group works with data centre operators at every stage of this lifecycle — from initial specification and equipment supply through to commissioning, load testing, and ongoing maintenance — delivering the infrastructure confidence that serious operators demand.
FREQUENTLY ASKED QUESTIONS
Power-related failures account for approximately 34% of all unplanned data centre outages — more than any other cause. This includes utility grid failures, UPS faults, generator failures to start, and power distribution issues. The vast majority are preventable with properly specified and maintained infrastructure.
Industry best practice is to maintain at least 24–48 hours of on-site fuel capacity, with supply agreements in place for extended refuelling during prolonged grid outages. Critical facilities — particularly those subject to extreme weather or grid instability — may hold 72–96 hours or more.
UPS systems provide short-term bridging power — typically 5 to 30 minutes — while the generator starts and synchronises. Generators then sustain the load for as long as fuel is available. Together they form a seamless transition: UPS bridges the gap, generator provides sustained backup.
Costs vary significantly by facility size, location, and existing infrastructure. What is consistent is that specifying N+1 redundancy at build is substantially less expensive than retrofitting it later. The Power Vault Group provides detailed cost modelling as part of the specification process.
Yes — generator paralleling, modular UPS expansion, and switchgear upgrades can all be delivered with careful planning to avoid live system disruption. The Power Vault Group has extensive experience in live-site upgrades where downtime is simply not acceptable.
Every generator, UPS, transfer switch, and PDU should report to a centralised monitoring platform in real time. Key metrics include load, temperature, battery state of health, fuel level, and fault codes. Proactive monitoring enables predictive maintenance — identifying degradation before it becomes failure.