1 min read
Migrating from Cumulus to SONiC for Broadcom ASICs
As a lot of people already know, Cumulus Networks was acquired by Nvidia/Mellanox in 2020, which was very exciting for some, but something of a...
3 min read
Marc Austin
:
Updated on July 29, 2026
Increase the utilization of an AI cluster by 15 to 20 percent, and the network has essentially paid for itself, often within a month or two. That economic reality sits at the heart of AI infrastructure, where networking is only about 20 percent of the capital spend but can gate the most valuable optimizations in the entire buildout.
I explored that math on the latest episode of The AI Hedge with Hasan Siraj, who leads strategy and product management for Broadcom’s core switching group and has partnered with Hedgehog for several years. Broadcom builds the silicon behind switches like Tomahawk, Jericho, and the Thor Ultra NIC, and it works closely with hyperscalers, enterprises, and the startup ecosystem.
Hasan and I examined how AI has changed the nature of the problem infrastructure teams are solving. Two decades of data center design centered on virtualization, running many applications on pooled compute. AI is different because the models do not fit on a handful of cores. Instead, tens or hundreds of thousands of GPUs and XPUs are needed.
“What binds all of this together, what glues all of this together, it’s the network. That’s why I say network is the computer for AI infrastructure,” he said. “It’s solving a fundamentally different problem. It’s a distributed computing problem as opposed to a virtualization problem.” That shift changes how organizations build, operate, and scale.
To make the complexity concrete, Hasan broke down the three domains of AI networking. Scale up connects XPUs directly inside a rack, where bandwidth, efficiency, reliability, and latency all matter because the traffic is essentially memory transactions. Scale out connects racks, where the priority is holding topology to two tiers before load balancing and congestion become unmanageable. Scale across links data centers, demanding a lossless fabric that can hold up across a hundred kilometers with line-rate encryption.
Each vector needs separate optimization, and the bandwidth numbers are climbing fast. Broadcom’s Tomahawk 6, the first 100-terabit switch in the industry, supports 64 ports of 1.6 terabits, with liquid-cooled and air-cooled designs to fit different facilities.
A persistent misperception, Hasan noted, is that buying an XPU from one vendor means buying the whole stack from that vendor. Hyperscalers never operated that way, taking the best XPU while building their own cost-conscious, best-of-breed network. “All of the other customers want to take this route. They want to go down the open route,” he noted.
In fact, the open route is not only good practice, in Hasan’s view, but it is the only viable path given the scale ahead. Hundreds of gigawatts and millions of XPUs will come online over the next four to five years, and no single vendor can supply all of it. That reality is why Ethernet is becoming the uncontested standard, and why the Ultra Ethernet Consortium has worked to replace proprietary silos with high-performance open standards.
The consortium’s value shows up in specifics. The RDMA protocol underneath these clusters is more than two decades old and was never designed for multipathing, out-of-order data placement, or selective retransmits. The consortium’s 1.0 specification modernizes it, and Broadcom has implemented those features in its Thor Ultra NIC, giving customers consistency across the board. When equipment inevitably comes from multiple vendors running different operating systems, Ethernet stays the one constant, and open APIs let a partner like Hedgehog present a single, consistent view across all of it.
Asked what teams financing and building clusters should fear most, Hasan did not hesitate. “The number one risk is losing your architectural sovereignty,” he said. Proprietary stacks may offer a sliver of short-term optimization, but the trade is the strategic flexibility a business needs in an unpredictable future. Lose control of the architecture, and the entire business strategy becomes tethered to another company’s pricing and constraints. His message: own your architecture, own your network, own your future.
The thread running through the conversation is that the AI buildout is too large for any one company to own end to end. With hundreds of gigawatts and millions of XPUs arriving over the next few years, interoperability stops being a preference and becomes the only way the industry can move at the required speed. Open standards are what let operators mix vendors, tune for their own workloads, avoid rebuilding from scratch each time a new generation of silicon lands, and own their destiny.
1 min read
As a lot of people already know, Cumulus Networks was acquired by Nvidia/Mellanox in 2020, which was very exciting for some, but something of a...
1 min read
The last couple of weeks have been filled with news of big-ticket AI acquisitions. It was against this background that I spoke with Ashmeet Sidana,...
1 min read
NANOG is the North American Network Operator's Group. It's a non-profit association where network operations professionals meet to share knowledge....