Nvidia’s LPX racks for AI inference accelerators have entered full production, the company has confirmed.

Unveiled at GTC back in March, the rack-scale platform came about following Nvidia’s acqui-hire of the eponymous startup. LPX is liquid-cooled and houses some 256 Groq 3 language processing units (LPUs) interconnected through some 640 terabits per second (Tbps) of scale-up bandwidth.

Nvidia LPX LPU rack
A glimpse inside the LPX from GTC 2026 – Sebastian Moss

Inside the LPX rack itself are BlueField-4 data processing units, Vera CPU racks, and STX storage servers all tied together with the recently debuted Spectrum-6 Ethernet networking tech.

The platform is not a replacement for Nvidia’s flagship NVL72 platform, but rather a complementary add-on for operators wanting to power ultra-low-latency AI inference workloads.

During the Hot Chips event in Palo Alto this week, Nvidia cited industry benchmarking results that saw the LPX platform support 3,400 output tokens per second. The chip giant claims it can be used to drastically reduce the time it takes to perform agentic-related tasks from hours to mere minutes, offering 4x faster responsiveness for agents and latency-sensitive workloads.

Nvidia CEO Jensen Huang said the LPX will “transform how intelligence is produced, delivering another giant leap in AI throughput, efficiency and responsiveness.”

“Inference is the growth engine of AI. Nvidia Grace Blackwell and NVL72 revolutionized large language model inference with an unprecedented leap in performance and efficiency,” Huang said. “Vera Rubin extends that vision with workload-optimized AI factory configurations designed for the era of agentic AI, advancing the performance frontier with LPX for ultrafast token generation.”

The LPX platform is expected to be available later this year.

Among its early adopters is neocloud darling Nebius, which plans to bring the Groq 3 LPX platform into its Token Factory offering. Danila Shtan, Nebius’ chief technology officer, said the move will make “every step of an agent’s loop feel instant.”

Groq, the chip and cloud company Nvidia's LPU tech has been licensed to, said it will also be one of the first companies to offer access to the new hardware.

Sinclair Schuller, Groq CTO, said: “Our customers expect Groq to be at the forefront of AI performance, and this platform represents a major advance for next-generation workloads. Together with Dell Technologies, we’re excited to deploy Nvidia Groq 3 LPX and Vera Rubin NVL72 at scale and make this capability broadly available.”