Lambda, the Microsoft-backed neocloud, has given the world a glimpse at Nvidia’s upcoming co-packaged optics (CPO) networking switches after it was given early access.
With the hardware in “full production” ahead of an anticipated release later this year, Lambda was given an early look to try it out for use in its GB300 NVL72 systems.
The neocloud was given a Quantum-X Q3450-LD switch – the InfiniBand-based offering – which comes in a 4U form factor, housing 144 x 800G ports. It offers up to 115.2 terabits per second (Tbps) of non-blocking switching capacity. The switch is liquid-cooled, with four UDQ4 liquid cooling connections via dual internal loops.
Nvidia’s embrace of silicon photonics sees it look to boost bandwidth demands and improve connectivity between servers in mammoth GPU clusters.
Among the early advantages Lambda’s engineering team touted was the Q3450-LD switch was its potential power saving capabilities. CPO switches do not require traditional pluggable transceivers, as the optical engine is already placed alongside the four ASICs inside the system.
Instead, Nvidia’s switch features fiber-array connections, with 18 removable light-source modules feeding the MPO (Multi-Fiber Push-On) ports. That means the electrical path between the optics and the switch ASIC is drastically reduced, and helps to boost latency while reducing potential signal losses.
Rich Underwood, HPC systems architect at Lambda, explained that a traditional 72-port switch would require 72 transceivers, each of which requires 25 watts of power.
“That equals a lot of power savings. So all of that power saving equals more efficient GPU and compute time, and more compute time equals more tokens for us,” he added.
“The benefit of CPO is that it minimizes the electrical channel,” explained Ashkan Seyedi, director for networking at Nvidia. “Now, the electrical channel that leaves the host turns into the optical domain immediately on the package, as opposed to traversing with the faceplate.”
“The reduction of lasers, not requiring digital signal processors (DSPs), and simply relying on the house fibers to be clean and ready to go, we’re able to stand up the system faster,” Seyedi added.
Beyond the concept of less power means more GPUs to play with, the folks at Lambda contend the switch to CPO could help improve network reliability, with the removal of transceivers taking out a potential failure point.
“When we’re talking about Superintelligence Cloud, we’re looking at a 128k GPU cluster, we’d potentially need five million lasers,” Underwood explained. “In CPO, that drops us down to one million. That total saving equals power reduction, which equals more GPUs, more tokens.”
TSMC is hard at work manufacturing the co-packaged optics iterations of both Quantum-X and the Ethernet-based Spectrum-X Ethernet switches, with the foundry having worked closely with Nvidia to create photonic integrated circuits.
Despite its adoption of CPO, Nvidia will still be making use of copper, with CEO Jensen Huang confirming during his GTC keynote that copper would still have its place as part of the company’s interconnect solutions, including for its next-generation Feynman architecture, having previously said the firm would “stay with copper for as long as we can.”
Nvidia networking SVP Gilad Shainer would later tell SDxCentral that using both technologies was about “doing the right design.”
“There is nothing better than copper," Shainer said. "Copper is zero power. Copper is passive. There's no active component. … If you can use copper, use copper. And if we can use copper, we want to use copper.”
Comments