“We utterly reject the notion that one company could have a monopoly on AI or AI innovation,” Forrest Norrod, EVP and GM of AMD’s data center solutions group said last week.
That one company is, of course, Nvidia, currently valued at $3.5 trillion and the world’s second most valuable organization at the time of publishing.
While AMD isn’t exactly struggling - itself valued at $197 billion - the company seems to have realized that trying to out-Nvidia Nvidia is not a viable business plan, so has instead decided to go all in on promoting an open standards approach to AI innovation and win over customers who may have grown weary of Nvidia’s walled garden.
During her keynote speech at the company’s Advancing AI conference in San Jose, California last week, CEO Dr. Lisa Su said that AMD is “the only company committed to openness across hardware, software, and solutions,” adding that for both the chipmaker, and the industry at large, “openness shouldn't be just a buzzword.”
AMD recently closed its acquisition of hyperscale server maker ZT Systems, in addition to acquiring silicon photonics startup Enosemi, AI software optimization company Brium, and will bring the hardware and software engineers from recently shuttered Untether AI into the fold.
“Over the last few years, over the last year, we've actually done more than 25 strategic investments that have been a great way for us to build new relationships and also support the AI, software, and hardware leaders of tomorrow,” Su said during her speech.
Banding together
Networking is one such area where AMD has fully committed to an open ecosystem approach, primarily through its work with both the Ultra Ethernet Consortium (UEC) and Ultra Accelerator Link (UALink), the former of which AMD is a founding member.
As Soni Jiandani, VP of AMD’s networking and technology solutions group, explained, when it comes to AI networking, “the only two options... that customers have at their disposal today is either InfiniBand, which doesn't scale, or ethernet, which scales but was not designed to run AI networks.”
With regards to scale-out networks, Jiandani said that, through the work carried out by the UEC, those involved with the open standard project are now able to scale their network 20 times more, when compared to InfiniBand, and 10 times more when compared to classic Ethernet.
The UEC released its 1.0 specification last week, and AMD announced Oracle Cloud Infrastructure would be deploying what it claims to be the first Ultra Ethernet-compliant NIC, the Pensando Pollara 400. Designed to support scale-out environments containing up to one million GPUs, AMD says that its Pollara 400 networking card offers a 10 percent higher RDMA performance compared to Nvidia's CX7 and can improve RDMA performance by 25 percent compared to traditional RoCEv2.
Additionally, Jiandani said the work being done by the UALink consortium also promises to deliver networks twice the amount of scale compared to Nvidia's NVLink.
“Unlike NVLink from Nvidia, UALink being totally open means our customers can use any GPU, any CPU, and any switch for their scale-up architecture,” she said. “From a scale perspective, the UALink 1.0 supports the ability to connect up to 1000 GPUs together, and that happens to be twice the scale of Nvidia's NVLink.“
However, it’s worth noting that while the limitations of InfiniBand and NVLink were discussed multiple times throughout the week, no mention was made of Nvidia’s own Ethernet offering – Spectrum-X Ethernet – which is currently being used to support xAI’s Colossus supercomputer in Memphis, Tennessee.
Collaboration is King
Collaboration was certainly the word of the day at Advancing AI, with the company using the conference to highlight its partner ecosystem, with Meta, xAI, Oracle, Microsoft, Cohere, Humain, Red Hat, Astera Labs, Marvell, and surprise guest of honor, OpenAI CEO Sam Altman, all trooping across the stage during the three-hour long keynote to discuss how they are deploying AMD hardware to support AI workloads.
Furthermore, in a briefing after the keynote, Norrod said that AMD’s newly announced double-wide Helios rack started life as a specific design for two hyperscale customers, influenced directly by their requirements.
However, AMD does not have a monopoly on acquisitions or partnerships, and you’d be hard pressed to find a company collaborating with AMD that isn’t also working with and deploying Nvidia GPUs to support their AI workloads – although, when everyone is an Nvidia preferred partner, is anyone really a preferred partner?
Saudi Arabia’s AI venture Humain, for example, spoke on stage with Su about the 500MW compute deal with AMD it signed last month, but the newly formed company is also currently waiting on 18,000 Nvidia GB300 to support a separate 500MW data center deal.
Oracle announced plans to deploy an AI cluster with up to 131,072 of AMD's new MI355X GPUs, but has also this week made its Nvidia GB200 NVL72 OCI Supercluster with 131,072 Blackwell GPUs generally available.
And OpenAI said it will deploy AMD's upcoming MI350 accelerators as part of its growing compute portfolio, but is simultaneously one of the largest users of Nvidia GPUs, primarily bought by its compute providers Microsoft, Oracle, and CoreWeave.
Furthermore, most hyperscalers are already developing their own chips – largely to reduce their reliance on Nvidia GPUs – and despite AMD’s commitment to openness, it’s unlikely that the chipmaker will be sharing any trade secrets with its customers anytime soon.
While having so many big names throw their support behind AMD and its hardware offerings is obviously of benefit to the chipmaker, without trying to be too disparaging of a clearly successful company that is building products that are in demand from customers, the question remains how much of that results from AMD being best in class, and how much is because they’re just not Nvidia?
It's been known that companies sometimes deploy AMD GPUs in an effort to get more leverage in negotiations with Nvidia, while AMD is also planning to significantly undercut Nvidia on price in an effort to gain market share over its rival.
Meanwhile, the difficulty of getting your hands on Nvidia GPUs – despite the company promising to deliver clusters in both the tens and hundreds of thousands – has forced customers to turn elsewhere or wait out massive lead times until supply can catch up to demand.
While current macroeconomic trends might have you thinking that truly anything could happen, Nvidia is not about to be toppled.
AMD knows this and instead of trying to outplay the chip giant at its own game, has instead decided to pursue the path trodden by many a politician trying to unseat an incumbent, offering what it believes is a credible alternative that will result in real progress.
And sure, the electorate is a fickle beast, but there’s no doubt that the increased demand for AI hardware can only be to AMD’s benefit. While Nvidia’s Jensen Huang may continue to be the darling of the tech conference circuit, promising GPUs, Oprah-style, to anyone who might want them, if last week’s event is anything to go by, those same customers would also really like quite a lot of AMD accelerators, too.
Comments