For many years, immersion cooling was a hard sell for hyperscalers: the new form factor and operational differences posed a unique and potentially unwanted challenge, especially as the previous rack densities meant immersion’s capabilities felt like overkill. Fears over what submerging valuable IT hardware in fluids might do to the kit have also been a barrier.
But with GPU power requirements growing fast, the data center industry is currently working furiously to try and figure out how to cool racks that could reach densities of 1MW. Hyperscalers are designing custom side pods for both power and cooling and, in the process, potentially re-architecting facilities from the ground up to cater for Nvidia’s latest hardware offerings.
Immersion cooling, a technology traditionally preferred by cryptominers rather than AI or cloud providers, might not currently offer 1MW tanks, but could provide operators with a way to reach higher densities without needing to rewrite the rulebook in quite the same way. Perhaps it is time to look again at immersion cooling, especially as some research suggests it can result in high uptime and more resilient hardware.
Immersion’s day in the sun?
For years, immersion cooling was the stronghold of the cryptominer. Crypto firms remain the biggest purveyors of immersion, but deployments hosting traditional IT are growing.
Placing a high-density immersion tank in a room requires less infrastructure than an air-cooled environment, and a growing number of Edge players are looking at offering immersion services within city metros, often in office buildings with spare capacity.
Immersion tank densities will vary, but it’s not unusual to see providers offering densities up to 100kW or more, with some reaching 200-300kW.
A number of cooling fluid providers – including Shell and Castrol – have their own immersion tank deployments, both to test their latest fluids, but also show potential customers how such deployments could work in the real world.
Established data center operators, such as Stack, Aligned, and CyrusOne, have said their latest liquid-cooled data center designs can support immersion cooling. NTT has an immersion cooling deployment at its Mumbai facility in India. Cloud firm PeaSoup has deployed immersion tanks in Centersquare data centers. Australia’s Firmus, another cloud provider, has an immersion deployment at an STT GDC facility in Singapore and is planning a large roll-out across Australia with NextDC.
A real-world view of immersion
One company that has been operating immersion cooling at a large scale for years is Australian firm DUG. Previously known as DownUnder GeoSolutions, the company was founded in Perth in 2003 to provide computing solutions for scientific data analysis and offers HPC as-a-service to customers.
DUG currently has HPC deployments in Houston, Texas; Kuala Lumpur, Malaysia; and Perth, Australia, totaling tens of thousands of immersion-cooled servers.
The company is also planning another Australian site in Geraldton, some 400km north of Perth in Western Australia. Clients include several universities, shipbuilder Austal, Australia’s CSIRO research agency, crypto firm HODL Ranch, and a number of oil and gas firms.
The company has been doing immersion cooling at scale for more than a decade, using mostly proprietary tanks and cooling systems, but doesn’t consider itself a data center or cooling firm.
“That’s not who we are,” says Matt Lamont, DUG managing director. “We predominantly make our revenue by processing and imaging seismic data.”
While it has historically kept its data center designs and technology to itself, DUG is starting to offer its wares to others. Cooling solutions provider Baltimore Aircoil Company, Inc. (BAC) has a licensing agreement with DUG to sell its patented immersion-cooling technology. DUG is also offering a new immersion-cooled data center module, known as a Nomad, to its customers who want to keep data hosted locally.
Lamont says the company designed its own tanks because it didn’t like what was available on the market at the time, and thought it could make a more reliable one for less. Its standard 26U tanks could originally go up to around 50kW – with the company pushing them to 60-70kW with some component tweaks – while the Nomad tanks can reach around 100kW thanks to the introduction of a second heat exchanger.
“All these things have been developed to help us in our numerical science business,” says Lamont. “Everything we have, compute-wise, is always immersed. We would never do it any other way. You’re nuts if you do it any other way."
“It just works”
The operational simplicity and resilience of immersion cooling means DUG can operate with very high uptime, without needing to invest in large amounts of redundant infrastructure that would be necessary in an air-cooled or direct-to-chip liquid setup.
While DUG’s seismic computing services are mostly in batch compute, the firm does offer some HPC-as-a-Service products, but doesn’t operate with a hyperscale-level SLAs for its customers – often offering something like 99.95 percent uptime. Yet Matt Lamont notes that any prolonged downtime would be “very serious” for the company.
Dan Lamont, DUG’s commercial officer at the time of our conversation and now acting CFO, notes that the company doesn’t operate in anything like a Tier III environment in its own data center.
While there is some backup power for storage and network systems, there’s no backup for compute. The company’s server fleet operates in a stateless configuration, so the individual servers failover to another machine.
“We don’t have N+1. We don’t have 2N. But we also run at around 99.999 (percent uptime) because of the resilience and simplicity of the system,” Dan says. While it has smoke detectors, the company doesn’t even have fire suppression in the data halls because the ignition temperature of its dielectric fluid is so high. The company notes that not all fluids have the same flashpoint, however.
“We had a faulty Nvidia GPU. We put it in and it shorted out and burnt up. It put off smoke, and it set off the alarms. But nothing happened,” says Matt. “We just switched off the power to that one GPU, and the rest of the tank kept running while we pulled it out. If that doesn’t cause any issues, nothing ever will.”
The tanks have dual pumps for redundancy, but very little else.
"When we started doing it, we’d keep spare heat exchangers,” adds Matt. “But we’d just never use them. But that’s our system, we worked on it for a couple of years to get it as simple as it is.”
DUG notably doesn’t have lids on its tanks, and the company says that occasionally dust has to be removed from the bottom – a clean-up process it claims is mostly aesthetic, as the dust doesn’t impact operations.
“A long time ago, we had a client come into the computer room with a cup of coffee, and he dropped half the cup into a tank,” Matt says. “And nothing stopped, everything kept running. We never did anything. It is really robust.”
Despite this, Matt notes the company has a track record for high uptime: “We ran a job in the not too distant past where we had 8,600 whole machines, running 38 days on a single job. That’s the sort of reliability we’re getting out of these machines, and a lot of it is based on the immersion cooling.”
Operational differences with immersion cooling
DUG suggests the training and skills needed to operate and maintain an immersion-cooled environment are essentially unchanged compared to air or direct-to-chip liquid-cooled environments.
Its proprietary tanks – developed by the company’s own engineers – were designed to be as simple and resilient as possible.
“They had good reason to make it robust and simple, because they were the guys that had to live with it,” says Matt Lamont.
The Lamonts say they don’t think they’ve ever had to top up or change the fluid in any tanks – and have done chemical analysis to check the fluid conditions, but have never seen notable changes. They joke that the only difference is that the engineers might get soft hands if they’re using a non-toxic immersion fluid without gloves.
However, immersion tanks in data halls do pose a different operational challenge to traditional racks. Servers and equipment must be lifted in and out of tanks vertically, rather than slid in and out horizontally.
In many deployments, it is common to have a portable crane – akin to the kind used to remove car engines – or a fixed overhead system if the room allows for it.
Once servers are removed, the dielectric fluid must be collected and cleaned off the removed servers – often requiring a drip tray and a cleaning solution. While not the most complicated process, it is a different way of operating that requires consideration.
Immersion cooling firm Submer is developing a robot to automate the installation and removal of servers from immersion tanks. The company first revealed the original version of its Autonomous Datacenter Assistant (ADA) back in October 2021, and previously said a new version would be out in late 2025. Submer isn’t the first to try and automate server deployments into immersion tanks.
TMGcore launched Otto, a two-phase immersion cooling system that used robotic arms to handle maintenance tasks around 2020. Rather than a mobile robot, however, Otto was a fixed system attached to the tank. Designed in partnership with Olympus Controls, it was equipped with an enclosure next to the tank housing backup servers and open slots for removed hardware.
For a long time, dunking servers in immersion fluid was done at the owner’s own risk. Making machines ready for immersion can take work if servers haven’t been specifically designed for it: fans need to be taken out (with a replicator or BIOS change installed instead), heat sinks and thermal paste removed. Cables connecting the servers to other equipment also have to be positioned correctly, or fluid may travel along them (a capillary action also known as wicking).
Until recently, none of the servers DUG bought had any kind of warranty beyond initial early life failures. This was primarily because few OEMs provided any kind of guarantees against immersed hardware. Several providers are now starting to offer warranties against immersion, and so DUG has been taking the companies up on this.
“We never needed it,” says Matt Lamont. “It was never an issue."
Immersion reliability: What do the numbers say
But it’s not just DUG that’s seeing high uptime for immersion environments.
During a talk at the Open Compute Project’s EMEA Summit in Ireland last year year, Austin Hipes, chief technologist, VP engineering at system integrator Unicom Engineering, provided reliability data from a large-scale immersion cooling deployment in Asia.
A total of 3,590 second- and third-generation immersion-cooled Xeon servers were deployed into a ‘Tier 1 data center in APJ’. The server units – from a “Tier 1 manufacturer you’d all recognize” – were deployed in a GRC ICEraQ FLEX series tanks with PAO4 dielectric fluid – a configuration Hipes described as “a very standard deployment.”
From Q1 2023 to Q1 2025, the second-generation Xeon servers saw a quarterly failure rate of just 0.9 percent, while the third-generation servers saw a 1.9 percent failure rate. When you remove early life failures (suggesting manufacturing defects) and cable issues, failure rates drop to 0.6 percent and 1.4 percent for second- and third-generation units, respectively.
“In both situations, taking out early life failures and cables or leaving them in, reliability is better than five-nines,” Hipes said.
Hipes noted that his company’s data on server failure includes service windows – suggesting even though “it can take a little longer to service an immersion-based server,” the overall uptime is still better. He notes that while there could be an extra half hour or hour of service time, you’re servicing less often because the overall reliability is better.
The facility in question also had air-cooled, cold plate-cooled, and immersion servers of the same model, and reliability data suggests immersion server downtime is around one-eighth that of comparative air-cooled servers. But the true numbers could actually be better.
“We believe these failure rates are actually much higher than what you’d see in the industry today and much higher than what we have done in other deployments,” Hipes said during the presentation. “There were quite a few things done during the conversion process that led to some reduced reliability on some of those components.”
The unnamed operator’s conversion work to make the servers ready for immersion was subpar compared to how a dedicated system integrator would operate, leading to more failures.
Conversion work was done on the data hall floor, potentially leading to debris; no torque drivers were used on the components; thermal paste was removed from the CPUs using standard tissue paper, leaving debris and generating static; heatsinks were mounted without proper insertion tools, causing misalignment and damage. The general rough handling of the hardware led to damaged motherboards and other components. DAC cable configuration also created extra stresses and more cable failures than normal.
“This was not an integration environment,” Hipes said. “Everything was just hand-torqued to what that technician thought it should be. The work was performed literally on data center floors, not even workbenches. There were a lot of handling issues reinserting CPUs after they were cleaned, and the heat sinks themselves were reapplied.”
One motherboard saw its CPU blow out because some loose conductive material wasn’t removed properly. The debris floated around the tank until it hit just the wrong point, landing right in between two different CPU voltage regulators and shorting them out.
“Even with this severity of damage, where the actual PCB was melting, there was still no impact to anything else in the tank,” Hipes noted. “They didn’t even know that this was shorting out; they just knew the server stopped working. All other platforms ran just fine.”
Despite what could be described as shoddy handiwork, the data center’s immersed servers still saw better uptime than their air-cooled counterparts.
“The reliability rates here, the failure rates are still actually better than air, even with these particular issues,” Hipes concluded. “[With] other deployments, which are typically not at this scale, the reliability is actually better. [This] is noticeably worse than other deployments we’ve been involved in.”
Comments