Microsoft has unveiled its Maia 200 inferencing chip.
Built on TSMC’s 3nm technology, the hyperscaler described the accelerator as an “AI inference powerhouse” that has been “engineered to dramatically improve the economics of AI token generation.”
Delivering around 10 petaflops of FP4 and five petaflops of FP8 compute performance within a 750W SoC (System-on-Chip) TDP (Thermal Design Power) envelope, the chip has a “redesigned memory system” comprising 216GB HBM3e at seven Tbps and 272MB of on-chip SRAM, alongside data movement engines. This, claims Microsoft, makes the Maia 200 the most performant chip from any hyperscaler, offering three times the FP4 performance of Amazon’s Tranium3 and Google’s seventh-generation TPU, Ironwood.
The chip uses a novel, two-tier scale-up network design built on standard Ethernet, and offers 2.8 Tbps of bidirectional scale-up bandwidth. Four Maia 200 accelerators can be connected within a tray with direct, non-switched links for “optimal inference efficiency.”
The same communication protocols are used for intra-rack and inter-rack networking via the Maia AI transport protocol to enable scaling across nodes, racks, and clusters.
The new Maia chip is the company’s most efficient inferencing accelerator, offering 30 percent better performance per dollar compared to its latest generation of hardware.
In a blog post detailing the chip, Scott Guthrie, EVP, cloud + AI, at Microsoft, said the Maia 200 was designed with data center availability in mind, and integrates with some of the company’s “most complex system elements,” including the backend network and second-generation, closed-loop, liquid cooling Heat Exchanger Unit.
Microsoft said it has already deployed the chip in its US Central data center region near Des Moines, Iowa, with the US West 3 data center region near Phoenix, Arizona, set to follow. Time from first silicon to first data center rack deployment was reduced to less than half that of comparable AI infrastructure programs, the company claims.
“Ultimately, customers want to be able to use AI to fundamentally change their business and transform. Increasingly, that’s not just around text prompts, it’s around multi-modal,” Guthrie said in a video discussing the chip. “This is the first Maia specifically designed with these large language models like OpenAI, as well as our own first-party models that we’re building in mind.”
He added: “Our investments in AI include, from the infrastructure side, both hyper-optimized data centers, we call token factories, specifically designed for AI processing machines, and we’re building and investing heavily in silicon solutions like Maia that allow us to optimize and yield the best token per watt per dollar.”
Comments