As data centers evolve into ever more complex infrastructures, they are increasingly beginning to resemble industrial facilities rather than their traditional role as digital storage hubs.

With power and thermal systems becoming tightly intertwined, effective heat management is becoming a critical design and operational priority. Although the exact pathways for innovation are still taking shape, one thing is certain: data centers are expanding their footprint as both major energy consumers and significant heat producers.

Thermal management in data centers starts at the chip level, where computation generates heat that must be carefully controlled through the entire system. This involves a chain of interconnected solutions, from liquid cooling within the racks and localized air management, to the transfer of heat through fluid loops, heat exchangers, and large-scale cooling infrastructure, such as plant rooms, external chillers, and direct expansion systems. Refrigerant-based products, including computer room air conditioners (CRAC units), remain a cornerstone of these processes. George Hannah, senior director of chilled water systems at Vertiv, explains his role in this evolution:

“Everything we do is about efficiently removing heat from the chip, getting it out of the facility, and either expelling it or reusing it. We are increasingly exploring heat recovery and reuse solutions, not just rejecting heat to the atmosphere, but capturing it for social good or for re-use by the local facility itself.”

This approach highlights how next-generation data centers are designed to maximize compute efficiency while managing heat responsibly.

Today’s liquid cooling needs

For decades, liquid cooling in data centers was often discussed as a future possibility – a ‘working towards’ plan aimed at economizing and reducing energy consumption. Over time, end customer priorities have evolved, and today, the focus is increasingly on maximizing compute performance, with new opportunities emerging to use AI for facility control, a trend largely unseen in the past.

“Recent advances in AI hardware, particularly from Nvidia, have really pushed liquid cooling into the now rather than further into the future,” says Hannah, adding:

“This transition has been driven less by efficiency improvements and more by necessity. As chip densities increased, traditional air cooling methods reached their practical limits. With this shift to liquid cooling has come consequential benefits that we are learning to leverage for the benefit of our customers and society.”

Chip densification is further influenced by new load patterns, such as AI workloads, which trigger internal services very differently from traditional applications. This creates fundamental changes in control loop systems: air cooling allows more thermal inertia and time to respond, whereas liquid cooling’s smaller thermal inertia requires rapid response and fluid circulation to manage heat spikes effectively.

Tyler Voigt, director of thermal controls at Vertiv, adds that from his perspective in thermal control management, these densification trends bring new requirements for system design:

“From a controls perspective, the feedback loop during high-density heat extraction has forced us to rethink event handling across the entire stack. Mechanically and from a controls standpoint, the way we design and respond to these systems has definitely changed over the past three to four years.”

Consequently, issues previously considered minor – like valve response times, physical actuation, and leak detection – have become critical.

The possibilities of power

The scale of modern data center deployments is putting tremendous strain on the power grid, which often cannot meet current demand – or the forecasted demand on the horizon. Hannah says:

“Many of the large hyperscalers or colos are trying to figure out how to solve that problem. Some partner with local power generation companies, buying future demand to provide coverage. Those that haven’t done so are trying to find alternative solutions.”

While a few global giants have the resources to buy or build their own nuclear power stations, this is not a viable option for most. The solution for many has become “bring your own power” (BYOP): deploying on-site power generation systems that serve the data center directly.

“When new data centers are designed today, instead of relying solely on the grid, they are integrating on-site power stations with their facilities. These on-site generators function like traditional power stations, and as heat engines, they produce substantial byproduct heat,” Hannah explains.

This high-grade, abundant heat opens new possibilities. Technologies such as absorption chillers, historically underutilized in data centers due to insufficient heat, can now be deployed effectively when coupled with BYOP systems. This flexibility extends to operational optimization as well.

“We’re seeing customers become more flexible with operating conditions if it helps them achieve goals like maximizing compute per watt or minimizing energy use. Controls are now smart enough to adjust automatically within defined limits to optimize performance and efficiency,” Voigt notes, adding:

“That shift is driving a complete evolution in control technology – from advanced software and algorithms to higher-capacity hardware at the unit level and more secure, connected systems across the stack.”

By combining on-site power generation, heat reuse, and intelligent controls, modern data centers can achieve greater efficiency, resilience, and operational flexibility than ever before. To manage and optimize this new level of system integration, operators are turning to digital twin technologies that can mirror real-world performance and enable data-driven decision-making.

The digital twin as a partner in performance

The digital twin methodology allows engineers to create theoretical models of systems to simulate responses and tune control algorithms accordingly. Operational or production-based digital twins extend this approach by using field and system data to continuously improve model accuracy over time.

“With digital twin work on the design side, we can tune algorithms to respond appropriately to different IT load inputs and step responses,” explains Voigt. “This lets us better understand equipment behavior so it reacts effectively under varying conditions.”

Working with Nvidia, Vertiv developed its SimReady 3D assets. The collaboration involved creating both the visual aspects of a digital twin and the back-end performance metadata compatible with the Nvidia Omniverse Blueprint for AI factories.

“We’ve been using these models to predict algorithm performance on the back end – a key part of our partnership,” says Voigt.

Voigt, who began his career as a mechanical engineer and has since moved into software and controls, brings end-to-end expertise from firmware and board design to supervisory and system-level controls. He has also helped develop industry standards around digital twins.

His advice to operators integrating complex thermal chains is clear: simplify system interfaces. Many existing setups add unnecessary complexity due to outdated or mismatched communication protocols.

“There’s a lot of complexity in how these systems communicate. On the lower level, different protocols and APIs perform very differently – some are more integrated and higher performing than others.”

Legacy building management protocols like Modbus and BACnet are now giving way to modern, high-performance, secure IoT protocols such as MQTT and AMQP. As systems grow more interconnected, standardized, capable communication methods are essential for performance, integration, and security.

Real-life implications of digital twins

With digital twins allowing engineers to validate system performance, responses to dynamic conditions, and interactions before deployment, they are becoming essential for validating complex power and thermal systems.

“Traditionally, customers tested each component – pumps, heat exchangers, power systems – individually in a factory acceptance test (FAT) and then again on site in a site acceptance test (SAT) to get everything working together,” explains Hannah.

“However, as systems grow larger and more integrated, with thermal and power systems interacting and technologies like absorption cooling connecting the two, physically testing full-scale setups is no longer practical. You can’t build a one-gigawatt test system in a lab.”

Voigt shares a practical example: on-site tuning had caused excessive valve repositioning, far beyond the valves’ designed lifecycle, reducing expected life to just 1.5 to two years. Using a digital twin, Vertiv virtually replicated the issue and tested 38 different control algorithm combinations, finding the optimal balance between minimizing valve movement and maintaining tight temperature control.

“As a result, we reduced testing time by 90 percent, cut temperature variance by up to 75 percent, and deployed the solution to production in just a day and a half,” he shares.

This example demonstrates how digital twins accelerate problem-solving and improve real-world performance.

Growing synergy across departments

The thermal chain and power train now operate less as separate systems and more as partners in a shared ecosystem, each dependent on the other for optimal performance. This growing synergy extends beyond technology, driving closer collaboration between traditionally separate teams across design, engineering, manufacturing, and operations.

“The growth is so incredible that customers are looking for products and systems they can deploy quickly – solutions that are easy to install, reliable, densified, cost-effective, and efficient,” says Hannah. “Right now, speed of deployment is the priority.”

This rapid expansion introduces new physical and logistical challenges. Factories and transport systems must now accommodate larger, more complex cooling assemblies, while design teams must consider how to produce, move, and integrate these expanded systems efficiently.

While data center densities continue to rise, proactive and predictive control strategies are essential. Larger equipment responds more slowly due to mechanical limitations, so relying solely on temperature or power feedback is no longer sufficient. Predictive visibility into IT loads allows thermal systems to anticipate heat generation, minimizing performance fluctuations.

“With loads increasing, reactive control isn’t enough – we need to predict when the heat is coming, not just respond once it’s already here,” says Voigt.

Above all, customers demand redundancy and reliability. Components across the thermal and power chains must operate in sync, responding seamlessly to failures or dynamic load changes.

“As more elements of this chain are deployed and power systems couple with thermal systems, they should be designed and balanced to work together. Our goal at Vertiv is to architect the whole loop to be as complementary and integrated as possible,” explains Hannah.

Service is a critical part of this end-to-end thermal chain – from liquid cooling to chillers and controls – supported by a global network of skilled service engineers. Deploying and commissioning utility-scale systems now needs expertise in controls, system logic, and integrated components – far beyond traditional hands-on skills. This depth of service reinforces system reliability and represents a key differentiator for Vertiv in an increasingly interconnected market.

Predict, optimize, scale

Amid technology advancement, new opportunities are emerging for predictive and feed-forward operations, where system operation and maintenance could eventually be automated through predictive modeling. While full automation remains a longer-term goal, steady progress is being made toward more intelligent, self-optimizing thermal management.

Hannah likens today’s AI boom to a modern space race, with global competition driving rapid innovation and reshaping the data center landscape in ways that often test the limits of traditional standards.

With data centers expanding from industrial- to utility-scale operations, they require greater design sophistication, advanced control architectures, and larger, more integrated thermal solutions. Amid these shifts, refining thermal management remains one of the most practical and high-impact levers for improving efficiency, reliability, and performance in the next generation of data centers.

Find out more about thermal management with Vertiv.

Explore more from Vertiv through its Power Innovation Day, here.