After another week in Vegas with Amazon Web Services (AWS), my ears are still gently humming with the echo of 60,000 voices through the conference halls.

As always, the AWS Re:Invent conference is big - not just in the number of attendees - and bold. And, as has been the case for the last couple of years now, emphatic on AI.

During the conference, one of the major news announcements as far as DCD is concerned was the launch of the latest generation of AWS’ Trainium chips - Trainium3, and teasers for the upcoming Trainium4. Revealed during Matt Garman’s opening keynote (which featured frankly traumatizing flashing lights), Trainium3 comes with a series of boasts - it’s the company’s first 3nm chip, and according to AWS can offer up to 4.4x more compute performance with four times greater energy efficiency.

Trainium looks to lift heavy weights

Available via UltraServers, customers can scale up to 144 Trainium3 chips with 362 FP8 petaflops and four times lower latency for training large models, and can even connect thousands of UltraServers to connect up to one million Trainium chips – around 10x as many as the previous generation.

DCD was awarded the opportunity to chat with Nafea Bshara, VP and distinguished engineer at AWS (and co-founder of Annapurna Labs, the company acquired by AWS for its chip offering), who explained that while the move to manufacturing the chip with 3nm added about “30 percent” more efficiency and compute, it was by no means the be-all and end all, telling DCD that the real gains came from their proprietary architecture offering.

As for the upcoming Tranium4 chips, Bshara was unwilling to share whether it would be developed using 2nm, but tells DCD: “We always go to the next and latest. The cutting edge, but not the bleeding edge. What I can tell you is that if you look at Trainium1 to 2, we had a four times gain, Trainium2 to three was three to five times. If we continue the mathematical series, it will be three or four to six times. We believe those big jumps are needed to drive costs down.”

“[Trainium4], it’s not going to be 3nm, it’s going to be more advanced, and it’s not going to be HBM3E,” Bshara says. He also noted - excitedly - that it will be compatible with NVLink - Nvidia’s scaling fabric.

It was interesting to hear CEO Matt Garman say during his keynote claiming that AWS’ Trainium2 chips (literally named for their purpose of training AI models) are also “the best systems in the world currently for inference,” with, of course, Trainium3 set to also benefit those workloads.

AWS has a custom chip that it makes specifically for inference workloads - Inferentia - though this statement suggests that Trainium is in fact superior even in Inferentia’s home grounds.

Bshara denied that this was the beginning phase of a move to drop Inferentia (though it remains somewhat telling to this writer that Inferentia barely got a mention throughout the conference, despite the majority of keynotes and talks focusing on AI Agents and the like).

According to Bshara, it comes down to the size of the model.

“We have all of them, just to be clear. It depends on the market and the use case. So, when we focused on small language models (SLM), we found that inference was a very different architecture from training. But when it comes to LLMs, we designed Trainium for training and found it was also better suited to inference in that situation,” he explains, adding that they are still used by customers.

Another notable highlight of the event was the launch of the AWS AI Factory offering, which will see the hyperscaler deploying its AI chips in other companies' data centers. This is a somewhat unprecedented move - with AI chips typically remaining in control of the hyperscalers behind them.

Bshara noted, however, that AWS has been practicing this to an extent for several years at this point, through its outposts offering, which features AWS’ Graviton chips.

According to him, the main motivation behind the offering is the increasing number of customers with sovereignty or residency requirements that still want to use AWS’ AI offerings.

While it is understood that the AI factory will be in customer data centers, the details on how the offering will be structured remain somewhat unclear. DCD is waiting for more information.

Trainium3 ultraservers rack
– AWS

AI moves from big investment to big “promises”

One thing that stood out to me was the shift in the AI story. While credence was given to AWS launching its latest AI chip and the company’s 3.8GW build out in the last 12 months, AWS seems to be steering the conversation away from the glitz of massive infrastructure investments and chip purchases and deployments, and starting to try to focus on the end goal - AI in practice.

Understandably, a lot of this was driven by an emphasis on “AI agents” - something that DCD does not include in its coverage - which saw at least two keynotes dedicated to the subject. The fundamental element that underlies this change in narrative, is really the shift from training workloads, to inference.

Day one of the conference also saw a panel dedicated to the subject of “physical AI,” again showing AWS’ desire to focus on the applications side of AI with start-ups discussing how they are using AI for things like robots, construction, manufacturing, and even to replicate five-finger dexterity.

During the panel, Nvidia’s head of robotics and Edge computing ecosystem, Amit Goel, noted that a partnership between the Edge and cloud computing providers will become increasingly important in supporting physical AI.

“Physical AI is different, and there are many problems that need to be solved,” he said. “The first is the data modality. The physical world has a lot of modalities involved. You have to understand force, you have to understand audio, etc. The amount of tokens or data generated is orders of magnitude more than with text data. What that means, from an infrastructure perspective, is that we need to build tools and computing platforms that will allow all of this data to be ingested for AI models to learn.”

That can, of course, be done in the cloud, but Goel notes that these applications are in the “physical world” and therefore “you have to have compute at the Edge.” Goel posits that instead we will see hybrid architectures, where things that need instant reactions will rely on Edge compute, and “something that can take a few seconds to reason about it can rely on a bigger model on the cloud.”

Regardless, the AI in use story remains somewhat nascent. During the Thursday keynote with Peter Desantis and Dave Brown, Desantis demonstrated some AI in action by displaying a real-time AI-generated animation on screen behind him, which transformed his visage into a black and white cartoon of Dr. Werner Vogels. While a fun stunt, the effect was somewhat hampered by the cartoon’s inability to move its mouth as Desantis spoke.

A final point of note - while not perhaps “newsworthy” was an announcement in the closing keynote by Dr. Vogels. Long known for his eccentricity and always fun talks, Dr. Vogels revealed that this would be his final keynote, telling the audience that it was time they heard from “young, fresh, new voices.”

Worry not, though, we were reassured multiple times that Vogels isn’t leaving Amazon behind. Perhaps he is just a bit tired out after 14 Re:Invents. I certainly am, after one action-packed Re:Invent.