Google is in talks with Californian chipmaker Marvell to develop two new inference chips.
Per reports from The Information and FundaAI, the discussions center around the development of a memory processing unit that would work alongside the hyperscaler’s Tensor Processing Unit (TPU) and a new TPU designed specifically for running AI models.
The proposed memory processors would divide AI workloads between themselves and the TPUs, depending on their compute and memory demands, two people familiar with the situation told The Information. The companies are aiming to finalize the design in 2027, with the aim of producing around two million units of the hardware, the report said, although it cautioned that discussions were still in their early stages and those figures could change.
The report went on to note that while Google has bought data center processors, including CXL controller chips, from Marvell before, those were off-the-shelf models, as opposed to the custom chips currently being discussed.
At present, Broadcom is the sole designer of Google’s TPUs, and earlier this month, the two companies announced they had entered into a Long Term Agreement that will see Broadcom develop Google’s processors until 2031.
However, the report said that Google has been considering alternatives to Broadcom since 2023, due in part to the high fees charged by the chipmaker that result in it receiving a percentage for every TPU produced – a figure that is spiraling as demand for the processors increases.
Google’s work on the development of new inference chips has reportedly been accelerated by the launch of Nvidia’s LPU at its GTC conference last month.
Elsewhere, Microsoft detailed its Maia 200 inference chip at the start of the year, and Meta recently unveiled the next four generations of its MTIA (Meta Training and Inference Accelerator) chips, which will primarily be used to support generative AI inferencing workloads. Last week, Meta announced it was partnering with Broadcom for the development of multiple generations of its MTIA hardware.
OpenAI is also developing its own custom inference chips in partnership with Broadcom, although, since the deal was announced in October 2025, no updates on the project have been provided by either company.
The growing popularity of inference-specific chips is being fueled by the shift away from AI training workloads alongside the rise of agentic AI tools and workloads requiring high-compute, low-latency performance, which general-purpose processors are unable to support.
Comments