Microsoft has brought its data center in Atlanta, Georgia, online.

The facility is the company's second 'Fairwater' site, its new data center design that primarily uses a closed-loop liquid cooling system. The data center is connected to the first Fairwater site in Wisconsin with an AI Wide Area Network using dedicated fiber optic cables.

Fairwater Atlanta
– Microsoft

Microsoft did not disclose the size or capacity of the data center - DCD has asked the company for further details.

The facility can support around 140kW per rack and 1,360kW per row, and includes hundreds of thousands of the latest Nvidia GB200 and GB300 GPUs.

Each rack houses up to 72 Blackwell GPUs, connected via NVLink. They are then connected in pods and cluster by a two-tier, ethernet-based backend network with 800 Gbps GPU-to-GPU connectivity. The company has its own operating system for our network switches SONiC, as well as a broad ethernet ecosystem, which the company says helps it avoid expensive vendor lock-in.

It has also worked with OpenAI, Nvidia, and others to define a custom networking protocol, Multi-Path Reliable Connected (MRC), to enable control and optimization of network routes.

The Atlanta Fairwater data center spans two stories, with Microsoft noting it made the decision specifically to reduce the distance between racks in three dimensions.

Thanks to Atlanta's reliable grid, Microsoft said that it was able to forgo on-site generation, UPS systems, and dual-corded distribution, allowing it to reduce time-to-market and operate at a lower cost.

As Microsoft brings more Fairwater facilities online, it plans to connect them all on the same network as a distributed supercomputer.

“This is about building a distributed network that can act as a virtual supercomputer for tackling the world’s biggest challenges in ways that you just could not do in a single facility,” said Alistair Speirs, Microsoft general manager focusing on Azure infrastructure.

“A traditional data center is designed to run millions of separate applications for multiple customers. The reason we call this an AI superfactory is [that] it’s running one complex job across millions of pieces of hardware. And it’s not just a single site training an AI model, it’s a network of sites supporting that one job.”