The escalating power density of IT equipment like Nvidia’s GPUs is transforming nearly every aspect of how data centers are designed, built, and operated. And nowhere is that transformation more dramatic than in cooling practices.
For decades, data centers relied on air cooling as the go-to strategy for managing heat emitted by servers and other equipment in these critical facilities. That approach was effective for managing power densities up to 30-40kW/rack – a threshold that the vast majority of hyperscale, cloud, colo, and enterprise facilities operated within. But today, power densities are dramatically surpassing the limitations of air cooling.
Nvidia’s roadmap illustrates that clearly. Nvidia’s current Grace/Blackwell GPUs have power density of 120-150kW+, depending on configuration. That is already well beyond the thermal management capabilities of air cooling, but upcoming GPU models have even higher densities. The Vera Rubin GPU coming this year will have a power density of 250kW+, the Feynman GPU’s density will be 400-600kW, and an unnamed model on the horizon will have power density of 1MW+.
For power density this high, air cooling is simply not enough. Thermal rejection needs to start at the chip with direct-to-chip (DTC) liquid cooling. But there is a problem: risk.
Introducing liquid into the data center dramatically escalates risk. In fact, it breaks a cardinal rule that has existed since the dawn of the data center industry: no liquid in the white space. It was such a sacrosanct rule that people entering the rack areas were required to leave their lattes and Diet Cokes at the door.
If the risk of spilled soda was a huge concern then, that risk is magnitudes higher with liquid cooling systems that are built into the fabric of the facility. Leaks in a liquid cooling system can rapidly cause catastrophic equipment failures and endanger operator safety.
Even brief interruptions to liquid cooling can lead to temperature spikes that cause costly equipment damage. And issues with cooling systems can bring data center operations to a halt for days or even weeks, which translates to lost revenue, SLA penalties, and reputational damage.
To mitigate those risks, organizations need to have the right strategy for DTC liquid cooling, including the right equipment, the right setup, the right operational model, and the right training. Achieving that would be easier if current operational models for traditional data centers translated well to these AI/HPC environments. But they don’t. DTC liquid cooling requires a fundamentally different approach with a very specific operational model and entirely new skills for the operations team.
Current operational methodologies and employee skills simply aren’t a match for AI/HPC. To prevent those expensive incidents above, organizations need a fundamentally new operational model, and their operations teams need entirely new skillsets. Those are critical to mitigating the escalated risks of liquid cooling and protecting the large investments that companies are making in AI.
What are the key ingredients for achieving that?
- Comprehensive assessment of operational requirements and data center design – Because no two DTC liquid cooling deployments are the same, a one-size-fits-all strategy will escalate rather than mitigate risks. The first step is a comprehensive assessment of facility design, customer demarc, technical specifications of the specific AI/HPC servers and GPUs being utilized, and the end-to-end cooling systems (including chiller, facility pumps, CDUs, tech loop and cold plats inside the servers). This provides the indispensable foundation for a customized operational strategy that mitigates risk.
- Effective commissioning and as-built mapping – Once a liquid cooling system is installed, it is critical to perform a comprehensive commissioning process to ensure that all cooling systems and other management systems are operating in ways that ensure the equipment is set up to mitigate risks. But this commissioning process also provides a blueprint for development of site-specific EOPs, MOPs, and SOPs that form the basis for an operational model that ensures that workers play a critical role in mitigating risks.
- End-to-end liquid cooling management and proper chemistry management – One of the responsibilities that will be entirely new for data center operations teams is the process of operating, managing and maintaining each element in the cooling management system, from the chiller all the way to the chip (i.e., cold plates) and back again. This requires classroom learning and hands-on training to understand not only key equipment like CDUs, thermal storage units, and drip pans, but also the specific protocols for the ways that they have been implemented in each facility.
- Salute offers an industry leading eLearning education and hands-on labs for our highly qualified staff to manage and operate your data center, ensuring their ability to support every aspect of the liquid cooling system. Another new responsibility for operations teams is chemistry management, which is critical to avoiding downtime incidents in DTC environments.
- Leak detection and incident response – When incidents do occur, it is critical for organizations to have processes and training that ensures that there is a rapid, effective response that prevents or limits the dangers to equipment and personnel. This can be the difference between a minor incident that can be addressed without an impact to operations and safety vs. an incident that escalates. To mitigate these risks, organizations need to have protocols and training that is based not on best practices today but that evolve continuously as equipment and best practices evolve.
This blueprint for DTC liquid cooling operations is critical for protecting the investments that our industry is making in AI/HPC. It is a blueprint based on extensive conversations and collaboration with companies like Nvidia, CDU manufacturers, chemistry manufacturers, chemistry distributors, leak management vendors, OEMs, hyperscalers, companies deploying AI workloads, and many other experts.
I am proud to say it is also the blueprint of Salute’s DTC Liquid Cooling service, which many of the leading companies in the AI market are utilizing for their data centers. Liquid cooling may be inherently riskier than air cooling, but the right strategy and execution can dramatically reduce risk.
More from Salute
-
Sponsored How do we build and retain a workforce that’s ready for what’s next?
The data center industry is growing at speed, but its greatest constraint isn’t technology – it’s people
-
Salute acquires data center engineering and consulting firm Northshore
Data center services provider to use Northshore's Seastack to improve lifecycle services
-
The UK’s AI ambitions demand a decarbonized data center strategy
The industry needs better data, clear reporting, and shared benchmarks
Comments