AI-driven demand continues to fuel the need for more reliable and resilient data centers. As data centers are now as mission-critical as aircraft, it is time to rethink the traditional redundancy architectures—and this requires moving beyond the static, over-provisioned approach.

Recent studies demonstrate new redundancy strategies to address these challenges. These include leveraging redundant data center systems to generate revenue by supporting the grid. Data centers are prosumers; they both consume and produce power. Similarly, shifting the cooling redundancy model from idle standby to active operation and optimising server workloads to improve efficiency have been demonstrated. This article reviews these recent advancements across three key data center pillars: cooling, power, and server-level redundancy, and argues that smarter implementation is more impactful than more hardware.

Cooling-level redundancy

The end goal of data center cooling is to keep the IT equipment operational without interruption. In a study cited by Cho et al. (2024), cooling system failures account for 51 percent of downtime, making it the most common cause of interruption in data centers. Hence, it is imperative to strike a balance between maintaining the reliability of data center cooling systems and reducing energy consumption.

Consequently, Cho et al. (2024) proposed a novel approach to cooling redundancy called the hot standby sparing (HSP). The authors compared HSP to the cold standby sparing (CSP) method. The CSP, a common cooling approach in data centers, keeps the redundant cooling systems in standby mode. The HSP approach, on the other hand, runs all the systems simultaneously, including the redundant ones, while operating at a lower load.

Cho et al. (2024) argued that the HSP-based cooling approach is more viable than the CSP method because HSP systems remain active, ensuring redundant systems can take on additional load during a failure. The authors demonstrated the effectiveness of employing variable-speed drive (VSD) and variable-frequency drive (VFD) technologies within the HSP redundancy model.

A 30MW data center, consisting of IT equipment consuming 8.8 kW per rack, was used for their case study. They reported a whopping 15 percent reduction in total cooling power consumption for the HSP-based system compared to the CSP system based on the same design capacity. Interestingly, a drop in cooling PUE from 1.23 (CSP) to 1.20 (HSP) was also noticeable in their case study. According to Cho et al. (2024), when combined with VSD/VFD and under N+1 or greater redundancy, the HSP-based cooling systems should be prioritised over CSP for maximum energy efficiency.

Power-level redundancy

Taking it a notch higher and thinking outside the box, Takci et al. (2025) proposed a novel approach for data centers to share power with the grid. The authors noted that redundant systems, such as UPSs, backup generators, and IT servers, could serve as a valuable source of backup energy for the grid during peak hours, rather than remaining idle.

A survey study cited by the authors reported that 10–50 percent of UPS capacity remains underutilised in data centers, even during power outages. They argued that in a situation where the IT load required is only 3MW, using an N topology for UPS power supply would be insufficient and render the system vulnerable during interruptions, justifying the importance of redundancy in data center operations. Higher-redundancy UPS topologies would enhance redundancy, reliability, and flexibility. For example, a 2N topology or 2(N+1) topology would meet the IT load requirements while providing the redundancy typically employed in high-availability data centers.

This high availability, combined with other redundant systems, such as backup generators integrated into the grid, offers a potential source of energy and flexibility that utility companies could explore in collaboration with data center operators. Grids continue to experience growing demand for power from electric vehicles, smart homes, and other technologies, making this collaboration financially rewarding.

Server-level redundancy

In addition to UPSs possessing underutilised capacity of up to 50 percent, Takci et al. (2025) reported that server utilisation typically ranges between 12–30 percent. Similarly, Shaukat et al. (2022) noted that an idle server consumes up to 66 percent as much energy as a server in use. Shaukat et al. (2022), while acknowledging the importance of redundancy in data center operations, further established that data center systems remain underutilised, with overall workload averaging 30 percent, and proposed an Energy-Aware Fault-Tolerant (EAFT) technique to improve data center energy efficiency.

The EAFT methodology aims to optimise energy efficiency without compromising SLA and performance. The authors delved into data center server-level operations, comparing the EAFT and the GreenCloud techniques, while simulating servers’ load at 75 percent, thereby increasing the number of active servers. Compared to GreenCloud, EAFT consumed more power but reduced unfinished server processing tasks by 80 percent during failures and required 27 percent fewer backup devices. The authors reported that one level of redundancy (eg. N+1) would be sufficient for data center workloads, as an additional level would only incur costs and overcomplicate the already complex data center systems.

Cross-layer comparison and insights

Evidently, the peer-reviewed papers summarised above collectively suggest that redundancy in data center operations is interdependent, and data center design and operation could be approached more strategically. For example, an optimised server-level redundancy would improve the cooling load profile of the cooling systems, subsequently impacting UPS/power requirements. Collectively, these studies provide a deep understanding of data center operational redundancy, highlighting the trade-off between efficiency and reliability in overall data center operations. Furthermore, from cooling-level to server-level redundancy, the reviewed literature established that the aforementioned strategies can be implemented in situ without redesigning or disrupting data center operations.

Cho et al. (2024) proposed a shift from the conventional cooling strategy (CSP) to a more energy-efficient one (HSP, combined with VSD/VFD), focusing on redundancy levels above N+1. Shaukat et al. (2022) showed that even at the server level, energy consumption could be more efficient. As a result, they proposed the EAFT technique for energy optimisation. Takci et al. (2025), on the other hand, have unlocked the potential of sharing untapped energy with the grid by leveraging data center redundant power systems (e.g., UPSs and generators).

Further work is needed to standardise a governance framework required to facilitate this revenue-focused symbiotic relationship between data center operators and utility companies. However, their work provides a critical roadmap for operators and providers looking to embrace this novel idea. Overall, these studies have shown that redundancy can impact data center operational energy efficiency and PUE. Additionally, unnecessary levels of redundancy would only increase operational costs, and N+1 is often a threshold for data center operations, depending on the Tier.

The key takeaways are:

  • How redundancy is implemented is more impactful than how much redundancy is in place. Essentially, a shift from static “N+1” configurations to a more dynamic, load-aware operation is needed. This shift can be strategically adopted in phases without requiring a data center redesign. The advent of AI will influence and facilitate this shift.
  • Data centers are prosumers—that is, they consume and produce energy because of their substantial energy loads and assets, from backup generators to the UPS batteries. Although there is potential to generate revenue with the right governance framework, otherwise underutilised or idle data center assets can serve the grid.
  • Reliability and efficiency of data center operations are not mutually exclusive. Implementing the right redundancy strategy can improve availability, reduce energy consumption and PUE. Ultimately, where N+1 is deemed sufficient and satisfies availability requirements, designing for N+5 or 2N would only add costs and complicate data center design.

Conclusion and way forward

The surge in artificial intelligence (AI)-driven, high-density loads drives the transition from static redundancy to dynamic, load-aware design. Hence, as AI workloads continue to grow, leveraging systems from already established industries offers a design strategy for data center designers. For example, incorporating aircraft system redundancy principles offers significant untapped potential.

The dissimilar redundancy principle is a common approach in aircraft design, in which different hardware and software are used to mitigate common-mode failures, adding an extra layer of protection for critical operations. Embracing these principles will further strengthen data center architecture against the unique challenges posed by AI-driven, high-density environments.