As organizations grow more dependent on digital technologies, expectations around service resilience and uptime have never been higher. These mission-critical environments support the systems that keep business running every day. But even the most advanced facilities can face risks that can impact performance. The strongest safeguard? Planned and reactive maintenance that is meticulously executed.

Today with AI and high-performance computing (HPC) often deployed side-by-side, ensuring consistent power and cooling for the GPU/IT infrastructure that support business-critical applications and services is essential. While often paired with AI, HPC workloads present their own unique challenges. Together, these workloads represent the most demanding test for modern facilities.

The 2024 annual outage analysis found that 54 percent of significant outages cost more than $100,000, and 16 percent exceeded $1 million. Rising costs are attributed to several factors, including labor, hardware replacement expenses, SLA penalties and longer recovery times. Increased dependency on digital services is the overarching reason – losing access, even for a few hours, significantly impacts the bottom line.

As data centers become more expensive to build and operate, uptime remains critical. Encouragingly, recent research from Uptime Institute suggests resilience is improving across the industry. Yet, before IT teams celebrate, it’s worth examining the details more closely.

GettyImages-1474257575
– Getty Images

Shift focus from individual components to system-level resilience

Power challenges remain a leading risk factor for facility performance, with uninterruptible power supplies, transfer switches, and generators ranking as the most critical points of potential failure. On the network side, performance can be disrupted by a wide range of issues, from switch and router faults to firmware bugs and even physical breaks.

IT systems and software come with their own set of problems, ranging from misconfiguration to component failure. Hardware and software may have become more reliable over the years, but component complexity is increasing faster than reliability. At the same time, refresh cycles are accelerating, causing the maintenance focus to shift from individual components to system-level resilience.

This is especially true when environments are running HPC and AI workloads, where downtime on just a single HPC cluster can derail a myriad of highly integrated business processes. With such complex architecture, a truly healthy data center – one that can withstand severe outages – is obsessively focused on reliability and availability. Reliability is about systems and individual components performing to their specifications, while availability is about what happens if there is an outage, where success is measured by time to recovery, the speed at which systems can be restored.

Ensuring nothing flies under the radar

Risk of outages can be mitigated and reduced by better architecture design and more developed redundancy plans, but they will never be entirely eliminated. Infrastructure monitoring tools are a key requirement, providing real-time alerts that prompt a first level response, but it’s what happens around them, the capabilities of the support team that will have a bigger bearing on time to recovery.

The way to continuously improve is to devise a proactive risk assessment process. Single points of failure should be identified and the reliability of key assets evaluated, from electrical and mechanical infrastructure to building management systems and physical security layers. For HPC clusters, this will mean accounting for power density and cooling demands that exceed those of traditional IT racks.

From power supplies and networks to racks and cooling, the way data center technology is maintained is more critical than ever. In pursuit of system-level resilience, the emphasis must be on elevating operational practices, which come down to two types of approaches: planned maintenance and reactive maintenance.

Executed properly, these strategies will also address the human factor that affects data centers. The proportion of human error-related outages rose by ten percent from 2024 to 2025, according to the same Uptime Institute survey, highlighting the importance of providing facility personnel with a strong foundation to implement procedures effectively.

Thorough maintenance plans, regular process reviews, and additional training can help build resilience in maintenance activities. In addition, obtaining certifications can support compliance, reinforce training, and improve overall process consistency.

In addition to training and supporting teams in best-practice procedures, you need to develop comprehensive maintenance plans that ensure nothing flies under the radar.

Planned maintenance with real-time visibility

All too often, maintenance processes are bogged down in poor vendor coordination and reactive decision making. A centralized system can house various documents and assets, creating a living record of maintenance history and essential data. Maintenance should be organized, effective and reliable.

A mission control center needs to be established and primed to expect the unexpected. Every third-party technician, service action and procedural step should be tracked from a single vantage point to make sure nothing is missed. Each technician must check in and out of the environment, providing real-time visibility into who’s onsite, what they’re doing and when they’ve finished.

Clear, well-structured procedural documentation and processes are essential for guiding support teams and ensuring alignment with operational standards. Just as important is the design of the preventative maintenance program itself. Maintenance activities should be deliberately grouped and scheduled to minimize repeated equipment isolation and reduce operational risk.

By coordinating related tasks into planned maintenance windows and allowing MEP systems time to stabilize between higher-risk activities, operators can maintain system resilience while improving efficiency. When maintenance activities are carefully planned and executed, they not only strengthen governance but also enhance the facility's overall performance.

Accurate and timely Field Service Documentation and Field Service Reports are also part of these procedural best practices, making sure every completed task is recorded and archived, creating a documented trail for audits, warranty tracking and root cause analysis.

Preventative maintenance on its own isn’t the only way to safeguard against failures. Predictive maintenance can be another layer of protection and resilience. By analyzing trends in asset performance, condition reports and failure history, operators can forecast which components are most likely to become an issue, and when.

Reactive maintenance with heightened vigilance

Even the best deployed maintenance schedules can reach their limits, as unplanned outages and critical system failures can still arise. The only way to combat the unpredictable is through heightened vigilance, where reactive maintenance plans are treated with the same rigor, visibility and coordination as preventative maintenance. 

Having an always-on 24x7 operations center is crucial, a single point of contact to execute a coordinated response to an issue, degradation in redundancy or outage. A typical event will surface through alert monitoring, triggering a process of seamless escalation. Resolution tracking procedures should follow, leading to timely remediation, perhaps involving the dispatch of vendor technicians under the terms of an SLA.

From first alert to issue resolution, every action must be traceable. For both planned and reactive maintenance, a transparent, auditable record of everything that happens becomes a valuable platform. Over time, this will help drive continuous improvement, incrementally reducing the odds and mitigating the risks of an unexpected incident taking down your data center.

Learn more about Ascent here.