Nvidia has announced the much-expected Language Processing Unit (LPU) chip, which came out of its semi-acquisition of chip designer Groq.

The Nvidia Groq 3 LPU will be available in liquid-cooled LPX racks, which feature 256 LPUs with 128GB of on-chip SRAM and 640 TBps of scale-up bandwidth. The rack is focused on low-latency AI inference workloads.

On Christmas Eve, Nvidia spent $20bn to license Groq's tech and hired its leadership team, including CEO Jonathan Ross, as well as other staff.

Nvidia Logo
– Sebastian Moss

"Recently, Nvidia licensed the Groq IP, and it's interesting to contrast these two kinds of processors, GPUs, with their large memory and amazing floating point performance and high throughput and token rate," Nvidia data center head Ian Buck said.

"But this other processor, the LPU, is optimized strictly for that extreme low latency token generation, offering token rates in the thousands of tokens per second. The trade-off, of course, is that you need many chips in order to get that kind of performance. And the economics, or the tokens per second per chip, is actually quite low."

The company aims to combine both chip approaches "to get to that multi-agent future," Buck said.

"These two processors will combine the extreme flops of GPUs and the bandwidth of LPUs into one. Let's contrast them: A GPU with its 288 gigabytes of memory, compared to only 500 megabytes of stacked SRAM. The LPU is only one 500th of the amount of capacity per chip, but the bandwidth is exceptional, [with] 22 terabytes to 150 terabytes per second bandwidth."

The LPX rack will be available in the second half of this year, "coinciding with Vera Rubin."