Google has unveiled the eighth generation of its Tensor Processing Units (TPUs), consisting of two chips dedicated to AI training and inference workloads.

Dubbed the TPU 8t (for training) and the TPU 8i (for inference), Google said the hardware was designed in partnership with Google DeepMind and has “purpose-built architectures” to support model training, agent development, and inference workloads.

“Our eighth-generation TPUs are the culmination of more than a decade of development,” said Amin Vahdat, Google’s SVP and chief technologist for AI and infrastructure. “The key insight behind the original TPU design continues to hold today: by customizing and co-designing silicon with hardware, networking, and software, including model architecture and application requirements, we can deliver dramatically more power efficiency and absolute performance.”

The training TPU

In a blog post detailing the new TPUs, Vahdat described TPU 8t as a “training powerhouse” that has been built to “reduce the frontier model development cycle from months to weeks.”

A single TPU 8t superpod can scale to 9,600 chips, offering two petabytes of high-bandwidth memory (HBM) and double the interchip bandwidth of the previous generation, Ironwood. Google said the architecture delivers 121 exaflops of FP4 compute performance to overcome memory bandwidth bottlenecks, while still maintaining accuracy for large models, with the per-pod compute performance almost tripling when compared to Ironwood.

Google TPU 8t
Google's TPU 8t – Google

TPU 8t has 19.2Tbps of bidirectional scale-up bandwidth and 400Gbps of scale-out networking bandwidth, with Google introducing a new networking architecture to support the hardware. Called Virgo Network, the company said it supports a 4x increase in data center bandwidth and has been built on high-radix switches that reduce network layers.

Additionally, with JAX and Pathways, Google said it can now scale to more than 1 million TPU chips in a single training cluster, with Virgo Network able to link more than 134,00 TPU 8t chips with up to 47 petabits-per-second of non-blocking bi-sectional bandwidth in a single fabric. As a result, this fabric delivers more than 1.6 million exaflops with near-linear scaling performance, the company said.

Google has also introduced TPUDirect RDMA and TPU Direct Storage in TPU 8t. TPUDirect RDMA enables direct data transfers between the memory and network interface cards (NICs), bypassing the host CPU and DRAM to reduce latency. Meanwhile, TPU Direct Storage also bypasses the host CPU to enable direct memory access between the TPU and high-speed managed storage, “effectively doubling the bandwidth for massive data transfers,” the company claimed.

The inference TPU

The TPU8i, by comparison, has been built to handle the “intricate, collaborative, iterative work of many specialized agents” that are emerging with the advent of agentic AI.

Designed with more memory bandwidth to serve latency-sensitive inference workloads, the TPU 8i is scalable to 1,152 chips in a single pod. It delivers 11.6 exaflops of FP8 compute performance, with a total HBM capacity of 331.8TB per pod and 19.2Tbps of bidirectional scale-up bandwidth per chip.

Google TPU 8i
Google's TPU 8i – Google

Google said when it comes to the TPU 8i, it has also “redesigned the stack” to include four capabilities that eliminate the ‘waiting room’ effect – when user requests are intentionally queued or delayed to maximize hardware utilization.

These include pairing 288GB of HBM with 384MB of on-chip SRAM to stop processors from sitting idle; doubling the physical CPU hosts per server by moving to Google’s custom Axion Arm-based CPUs; doubling interconnect bandwidth for Mixture of Expert models; and reducing on-chip latency by up to 5x with the introduction of a new on-chip Collectives Acceleration Engine.

Consequently, Vahdat said that these innovations allow TPU 8i to deliver 80 percent better performance-per-dollar compared to Ironwood.

Both the TPU 8t and 8i run on Google’s Axion Arm-based CPU host and support liquid cooling technologies. The company said it has also optimized efficiency across the entire stack to deliver integrated power management that can adjust the power draw based on real-time demand, resulting in up to 2x better performance-per-watt compared to Ironwood.

“By owning the full stack, from Axion host to accelerator, we can optimize system-level energy efficiency in ways that simply cannot be achieved when the host and chip are designed independently,” Vahdat said.

Both chips will be generally available later this year and can be used as part of Google’s AI Hypercomputer – a cloud-based supercomputer architecture launched by the company in 2023 that combines performance-optimized hardware, open software, machine learning frameworks, and flexible consumption models.