Meta has announced the next four generations of its Meta Training and Inference Accelerator (MTIA) chip.
Dubbed the MTIA 300, 400, 450, and 500, Meta said the new chips have either already been deployed or are scheduled for deployment in the next 18 months, and will primarily be used to support generative AI inferencing workloads.
Each new chip will offer improvements in compute, memory bandwidth, and efficiency, Meta said. Furthermore, in a blog post detailing the chips, the company said that “given the rapid pace of AI innovation,” it has built the capability to ship a new chip roughly every six months.
A closer look
Already in production for R&R (ranking and recommendation) training, the MTIA 300 is described by Meta as a “cost-effective foundation” upon which the company has designed its subsequent chips optimized for generative AI workloads.
Made up of one compute chiplet, two network chiplets, and several HBM stacks, each compute chiplet is itself comprised of a grid of processing elements (PEs) which contains two RISC-V cores and a Dot Product Engine for matrix multiplication.
The chip has a TDP of 800W and offers 1.2 petaflops of FP8 compute performance, 6.1Tbps of HBM bandwidth, and 216GB of HBM capacity, alongside 1Tbps of scale-up networking and 200Gbps of scale-out.
The MTIA 400, by comparison, offers six petaflops of FP8 compute performance, with a TDP of 1200W. With regards to its HBM bandwidth, the chip provides a 51 percent increase to 9.2Tbps over its predecessor, with HBM capacity totaling 288GB.
For scale-up and scale-out networking, the MTIA 400 provides 1.2Tbps and 100Gbps, respectively. A rack with 72 MTIA 400 devices, connected via a switched backplane, forms a single scale-up domain and can support both air-assisted liquid cooling and liquid cooling technologies already in deployment in data centers, Meta said.
The MTIA 400 has already been tested in Meta’s labs, and the company is “on track” to deploy the chip in its data centers. The MTIA 450 and MTIA 500 are scheduled for mass deployment in early 2027, with both also providing 1.2Tbps of scale-up and 100Gbps scale-out networking.
At the system level, Meta said the MTIA 400, 450, and 500 all utilize the same chassis, rack, and network infrastructure, meaning each new chip generation can be dropped into existing data centers with ease.
Meta said the MTIA 450 has been further optimized to anticipate “the rapid growth in generative AI inference demand,” Meta said, noting that the chip is an advancement on its predecessor in four areas. These include the doubling of HBM bandwidth to 18.4Tbps, the introduction of hardware acceleration to make both attention and Feed-Forward Network (FFN) computation, and custom low-precision data type innovations.
Additionally, the MTIA 450 offers seven petaflops of FP8 compute performance, 288GB of HBM capacity, and has a TDP of 1,400W. The chip also supports mixed low-precision computation without incurring the software overhead associated with data type conversion, Meta said.
Finally, the MTIA 500 offers 10 petaflops of FP8 compute performance and has a TDP of 1700W and 384-512GB of HBM capacity. The chip uses a 2x2 configuration of smaller compute chiplets surrounded by several HBM stacks and two network chiplets, along with an SoC chiplet that provides PCIe connectivity to the host CPU and scale-out NICs.
It also contains the same hardware acceleration and data-type innovations as the MTIA 450 to address bottlenecks associated with inferencing workloads.
“While our large-scale production deployments of MTIA chips have demonstrated strong R&R inference capabilities, we expect the latest four generations — either recently launched or planned for launch in 2026 or 2027 — to push the boundaries of generative inference, enable R&R training, and lay the groundwork for future generative AI training,” Meta’s blog post, authored by Yee Jiun Song, Andrew Tulloch, Harikrishna Reddy, CQ Tang, and Vijay Thakkar, read.
“Each generation of MTIA has built on the lessons of the one before, is co-designed with our software stack, and is guided by the trajectory of future AI models… Together, they bring us closer to our goal to deliver today and tomorrow’s most powerful AI experiences to everyone on our platforms.“
Comments