We frequently talk about artificial intelligence (AI) and data center power consumption through a macro lens.
The statistics are undeniable, unavoidable, and frankly staggering. The International Energy Agency’s October report found that global data center electricity consumption would “increase significantly” by 2030, citing AI and digitization efforts as a driving force.
In December 2024, the US Department of Energy found that, while data centers consumed around 4.4 percent of the nation’s power in 2023, this could increase to 12 percent by 2028 - all while power generation itself is also growing.
AI is skyrocketing the power demands of data centers for two key reasons; the GPUs it requires have a high power consumption, and the training of AI models requires lots of these GPUs.
While Chinese startup DeepSeek claims to have trained its model using far less hardware than its rivals in the West, for most companies, AI development remains a resource-intensive pursuit.
But within the massive numbers, the gigawatt campuses, and the tens of thousands of GPUs per cluster, are the minutiae: the microseconds. The power consumption within the context of time.
AI workloads - generative AI and the training of large language models in particular - demand a profusion of power in just a fraction of a second, and this in itself brings some complications.
“When you are engaging a training model, you engage all of these GPUs simultaneously, and there’s a very quick rise to pretty much maximum power, and we are seeing that at a sub-second pace,” Ed Ansett, director at I3 Solutions, tells DCD.
“The problem is that you have, for example, 50MW of IT load that the utility is about to see, but it will see it very quickly, and the utility won’t be able to respond that quickly. It will cause frequency problems, and the utility will almost certainly disconnect the data center, so there needs to be a way of buffering those workloads.”
In December 2024, the Lawrence Berkeley National Laboratory released the 2024 United States Data Center Energy Usage Report which looked at the spikey power consumption of AI training workloads.
That report explains that while GPUs typically do the bulk of the computation for AI training workloads, they operate as part of complex nodes including “supervisory central processing units (CPUs), memory, and high-bandwidth interconnect.”
“This integration means that individual components can throttle others, creating characteristic load profiles,” the report says. “Empirical measurements consistently show that even in computationally intensive workloads, node-level power demand rarely approaches manufacturer-rated maximums.”
The report goes on to note that “the characteristic fluctuation in power demand for AI training as compared to more conventional HPC workloads... can be especially taxing to grid operators.”
There is some nuance to this, of course, depending on the utility and the region. It's a matter of scale - how big the data center is, and how much of the utility’s power draw it represents.
But real-world consequences are being seen in the grid. A recent report from Bloomberg, using data from Whisker Labs and DC Byte, found that massive AI workloads are leading to “bad harmonics” in the grid.
The report describes bad harmonics as a distortion caused by massive power use - comparing it to the static that can be heard when a speaker’s volume is too high. The result isn’t a smooth gradation of power use, but one with rapid rises and drops, which according to
Bloomberg, is actually impacting nearby residences. “The worse power quality gets, the more the risk increases,” the report says. “Sudden surges or sags in electrical supplies can lead to sparks and even home fires.”
According to Bloomberg, more than half of power quality-tracked households showing the worst distortions of power quality were within 20 miles of “significant data center activity.”
This activity could refer to one hyperscale facility or a cluster of data centers in one area.
Speaking to Bloomberg, Dominion Energy - which serves Northern Virginia among other regions, said that it had not noticed the same distortion levels that Whisker Labs’ data showed, but that “in a few rare instances, we have observed very brief periods of higher-than-normal harmonic disruption due to abnormal configurations or equipment issues when a new installation first comes online.”
Vertiv’s vice president of large power converters, Giovanni Zanei, conceded that it was a scale issue. “A few megawatts might not be such a huge problem,” he says. “A 100MW data center is certainly a different story.”
While there are some smaller AI deployments, the trend across the industry is for massive GPU clusters to be deployed by the hyperscalers and major tech companies.
In October 2024, Tesla said it would be launching a 50,000 GPU cluster in Texas, and planned a 300,000 cluster of Nvidia Blackwell B200 GPUs for the summer of 2025. Elon Musk’s AI firm xAI has a 200,000 GPU cluster dubbed Colossus, which it claims will ultimately be expanded to one million chips.
Meanwhile, Oracle offers a “supercluster” made up of more than 65,000 GPUs, and recently launched a “zettascale” supercluster with around 131,072 Blackwell GPUs. The likes of Meta, Amazon Web Services, Google, and Microsoft have also snapped up GPUs not by the handful, but by the truckful.
The power draw of individual GPUs is ever-growing, with the Blackwell range seeing the B200 GPU expected to draw up to 1,200W, while the GB200 - featuring two B200 GPUs paired with a Grace CPU - could reach 2,700W.
Powering up
When it comes to the rapid switching on of massive quantities of GPUs, i3’s Ansett is seeing customers turning to batteries to “buffer” the IT load, noting: “The batteries do the heavy lifting during the short term, but that load can then be transferred across to the utilities.”
Vertiv’s Zanei, however, does not posit batteries as an ideal solution for this problem, noting that within the rapid increase of power consumption lies the nuanced issue of voltage droop, or sag.
The vendor has been running tests on the issue of voltage droop, using its 1MW uninterruptable power supply (UPS) systems. “We tested this to see the scale and rate that it is happening, and it can easily be from, say, maybe 30 or 40 percent of the load, up to 150 percent of the load, in a matter of milliseconds, then stays there for maybe 10, 20 milliseconds, and goes back down,” Zanei says.
This complicates the matter even further. Not only is a sudden and massive demand for power being sent to the utility provider, but in less than a second it could disappear and then come back. It is the definition of “mixed signals.”
According to Zanei, the problem is not limited to the utility provider. That spasm is bad, regardless of the power source.
“Whenever a data center has to run on generators, the generators are not going to like this [voltage droop],” he says. “Or, in the case where a data center is running on primary power but not from the utility - for example, gas turbines. Again, the gas turbines are not going to like it, and you would need to potentially massively oversize the turbines to avoid an issue.”
Then there are batteries. Zanei says: “Batteries are meant for backup in which they are typically used every once in a while. But they aren’t designed for this kind of frequency - it could be hundreds or thousands of times per day.
“Even batteries with a higher cycle time are not made for that, and a warranty would certainly not be available.” Zanei adds that, if he were to estimate, he’d expect such taxing
workloads could halve the life of the battery. Zanei notes that having intelligence within the UPS to buffer or at least smooth over some of the power fluctuations is something that Vertiv is seeing customers do.
Vertiv is also advising customers on the use of UPS and batteries for voltage droop - regarding issues of the size necessary, type of canister, and whether having separate batteries and UPS systems for the data center and this exact purpose would be the best solution.
Alternatively, ultracapacitors could be the long-term answer, Zanei says. “The ultracapacitor is a kind of energy storage solution, and it's very good at moving power in and out,” he explains. “It stores less energy [than a battery], but it is excellent in cycling very quickly and could provide the power needed just for those milliseconds, or whatever the brief peak is, and then recharge during the valleys.”
Schneider Electric, on the other hand, has built power smoothing capabilities into its UPS.
Mustafa Demirkol, VP of data center systems at Schneider Electric, tells DCD that energy buffering solutions are already “incorporated into our topology,” adding: “We were designed for this before the AI fluctuations emerged, because we were seeing them in other applications as well.”
“When you look at AI load fluctuations, you need to look at this holistically. And it's not just about UPS, it is about starting from the server level. You may need to use some buffing and resources. Your UPS needs to be robust enough to handle these types of stress and load fluctuations, and, at the same time, you need to look at this entire system, not just the UPS but also at your switch gear. Everything needs to be able to handle this type of budget.”
Solutions at the chip level
Vertiv’s Zanei says his team is in regular conversations with chip companies. “When I talk to the chip guys, it’s clear that it’s a matter of making sure that the GPUs, or the servers overall, are at the maximum level. They need to run their GPUs and the data center itself in this way. And the way the load is drawn - the power taken up by their servers is all quite synchronous.”
DCD contacted several GPU manufacturers for insight into how they could engineer their chips to avoid this, and whether they are making an effort to resolve the problem. None were willing to comment on the matter.
Despite this, AWS and Annapurna Labs have made some moves with the second generation of their home-grown AI accelerator - Trainium. These chips differ from GPUs in both an architectural standpoint, and their end capabilities.
“If you look at GPU architecture, it's thousands of small tensor cores, small CPUs that are all running in parallel. Here, the architecture is called a systolic array, which is a completely different architecture,” says Gadi Hutt, director of product and customer engineering at Annapurna Labs. “Basically data flows through the logic of the systolic array that then does the efficient linear algebra acceleration.”
A similar approach has been adopted by other chip designers, such as Furiosa AI.
Trainium was designed for AI training workloads, and does not have the more general purpose computing capabilities that are seen with GPUs. “It cannot do graphics or weather forecasting, or all sorts of things that GPUs can do really well,” Hutt says. “But these chips are designed to do linear algebra, accelerate linear algebra, and pace in high utilization.”
Peter DeSantis, SVP of utility computing at AWS, discussed during his keynote at AWS Re:Invent 2024 titled the complexity of AI scaling laws: the physics behind maximizing the potential of GPUs - or other accelerators - for AI workloads.
Citing a 2020 paper: Scaling Laws for Neural Language Models, DeSantis described AI workloads as “scale up,” not “scale out.” The scaling laws hypothesized that model capabilities improve as you scale up various factors, including the number of parameters, the data set, size, and amount of compute.
As a result, there has been a consistent and considerate push towards bigger and more compute-intensive models which have indeed become more capable, with the aim of bringing as much compute and high-speed memory into the smallest space possible, wired together with low-latency connectivity.
The “smallest space possible,” is perhaps the key element.
Smaller spaces means closer together, thus with shorter cables and lower latency connections - as DeSantis explains: “If you have things closer together, you can use shorter wires to transmit data between them, which means you can pack in more wires. It also means you have lower latency and you can use more efficient protocols to exchange data. So this seems simple enough, but it's a very interesting challenge.”
AWS’ own chips require a lot of power too, and with that comes the problem of voltage droop. One of the ways the Trainium2 chip has been adjusted to avoid this is through the use of relatively “big” wires - power vias - which enables them to move power at low voltages.
“A semiconductor uses the presence or absence of tiny electrical charges to store and process information. So when chips encounter voltage droops or sags, they typically need to wait until the power delivery system adjusts, and waiting is not something you want to do with a chip,” says DeSantis.
It is actually far more efficient to move power at higher voltages in a data center, so the power is progressively stepped down as it gets closer to the chips - the final stage of which is before it enters the package.
According to DeSantis, this is typically done via voltage regulators positioned as close to the package as possible. In an effort to reduce voltage droop when compared to the Trainium1 chip, Annapurna Labs has brought the voltage regulators even closer to the chip - actually under the perimeter of the package.
“Doing this is quite challenging, because the voltage generators generate heat, and so you have to do some novel engineering. But by moving those voltage regulators closer to the chip, we actually can use shorter wires, and shorter wires mean less voltage droop.”
AWS may have been able to resolve this issue at a chip-level for its Trainium2 chips, but ultimately the very design of GPUs renders it unrealistic that we will see a chip-level resolution without them fundamentally revolutionizing their approach.
And for the most part, that isn’t going to be the priority. Yes, they have built-in “spikiness,” but GPUs are also incredibly powerful, and popular. The incentive to reinvent is limited.
In the meantime, data center operators will have to turn to their UPS in the hope of smoothing the power profile of AI training workloads.
If they don’t, the consequences to the grid could be severe. Data center operators could be stung, too.
As Schneider Electric’s Demirkol points out: “There probably will be failures across the upstream level, and you could even damage your GPUs.”
Comments