1 min read
The Enterprise AI Journey: Five Phases from ChatGPT to Owning Your AI Factory
Every enterprise we talk to is somewhere on the same road. The faces in the boardroom change, the industry changes, the model du jour changes — but...
4 min read
Marc Austin
:
Updated on August 11, 2026
AI infrastructure no longer just scales by adding accelerators. It scales by how well those accelerators communicate. Once GPUs reached sufficient performance, the limiting factor moved to the network that coordinates them — and today the network, not compute, determines the utilization, reliability, and return on every AI dollar you spend.
The industry has caught up to this reality. Gartner's recent research on hyperscaler AI networks reaches the same conclusion, projecting that by the end of 2027 more than half of all data center switching spend will support AI workloads, and advising enterprise infrastructure leaders to adopt hyperscaler networking practices selectively rather than wholesale.
That advice is right. But most of the conversation about "learning from hyperscalers" fixates on the artifacts — the topologies, the congestion control schemes, the telemetry pipelines. Those matter, and we'll get to them. The deeper lesson is about how hyperscalers buy and operate infrastructure: they never let a single supplier, a single chip, or a single architecture hold their roadmap hostage. Their real advantage isn't scale. It's optionality.
Look at how AWS, Google, Microsoft, and Meta actually build AI fabrics and a pattern emerges that has little to do with any one vendor's product line. They standardized on Ethernet for scale-out networking — not because Ethernet won a spec-sheet shootout, but because Ethernet gave them a multi-vendor ecosystem: merchant switching silicon from competing suppliers, multiple switch manufacturers building to open designs, interchangeable optics, and network operating systems they control rather than license.
Compare that with the proprietary alternative, where the interconnect, the switches, the optics, and the software all come from one vendor. Every refresh cycle, every capacity expansion, and every price negotiation happens on that vendor's terms and that vendor's timeline.
Ethernet's emergence as the default AI fabric — now conventional wisdom across the industry and echoed in analyst guidance to standardize on Ethernet with RDMA (RoCE) — is really the victory of supplier diversity. The protocol was never the point. The ecosystem was.
The AI hardware market moves faster than any infrastructure market in memory. Accelerator generations arrive roughly annually, from NVIDIA, AMD, and a growing field of custom silicon - Sambanova, Cerebras, Etched and several others. Switch port speeds have jumped from 400G to 800G, with 1.6T on the horizon. Optics, power envelopes, and rack densities shift with every cycle. Whatever you deploy today will be mid-life in eighteen months.
In that environment, an open, multi-vendor fabric is a speed advantage:
You adopt each new generation when the market ships it, not when your incumbent gets around to it. When any qualified supplier delivers 800G switching or next-generation optics, you can deploy it — no waiting for a single vendor's roadmap to catch up.
You mix accelerator vendors without re-architecting. An Ethernet scale-out fabric doesn't care whose GPUs sit at the edge of it. As new accelerators earn a place in your clusters, the network absorbs them.
You hedge the supply chain. Lead times on GPUs, switches, and optics have made "second source" a first-order design requirement. Multi-vendor procurement means a constrained supplier is an inconvenience, not an outage in your buildout plan.
Hyperscalers refresh their fabrics on their own schedule because no single supplier can gate them. That is a discipline any enterprise can copy — and in a market that reinvents itself every twelve months, it may be the single most valuable one.
The same optionality that buys speed also buys leverage. Merchant silicon, open network operating systems, and multiple hardware suppliers create genuine price competition on every refresh — on switches, on optics, on support renewals. Lock-in premiums quietly compound across all three; supplier diversity strips them out.
The economics extend beyond procurement. Consider oversubscription: fully nonblocking fabrics across large clusters are expensive, and for most enterprise AI workloads, unnecessary. Treating oversubscription as a cost-performance dial you deliberately set — rather than defaulting to hyperscale-grade nonblocking designs — is where much of the savings lives. This is a place to emulate hyperscaler judgment, not hyperscaler specifications.
And then there's the number that dominates every other line item: GPU utilization. A GPU cluster is the most expensive asset in your data center, and the network decides how much of it you actually use. Industry analyses of traditional Ethernet fabrics under AI workloads consistently show effective bandwidth well below line rate once congestion and poor load balancing take their toll — roughly 60% is a common finding. Modern congestion control and adaptive routing can push that to 95%. Our own analysis puts the difference at a minimum of $50,000 per GPU per year in recovered value. Across a cluster of hundreds or thousands of accelerators, network design stops being an engineering detail and becomes a board-level economic decision.
For the conversation with your CFO, frame it this way: in a market this volatile, a single-vendor commitment is an unhedged bet. Optionality has quantifiable value, and it accrues every refresh cycle.
A word of caution: hardware choice without automation just multiplies operational pain. Nobody wants three vendors' worth of CLI dialects.
Hyperscalers pair supplier diversity with a software layer that makes heterogeneous hardware operate as one system. Configuration is declarative — operators describe intent, and the system converges on it. Changes flow through automation, not console sessions. Telemetry spans the entire fleet, correlating network behavior with workload performance, because you cannot improve what you cannot see. Useful work completed between failures — goodput — is the metric that matters, and it is only measurable with end-to-end observability across every vendor's gear.
This is the piece enterprises most often miss. The unit of adoption isn't a switch. It's an operating model. Adopt the hyperscalers' hardware flexibility without their operational software discipline and you inherit their complexity without their advantage.
Pulling it all together. This is what is worth adopting for an enterprise:
Conversely, business should avoid:
This is the philosophy Hedgehog was built on. Hedgehog delivers AI-native network software with the hyperscaler operational model built in — Ethernet scale-out with modern congestion control and adaptive routing, VPC-style multi-tenancy with limited blast radius, intent-based automation that takes clusters from zero to inference in hours, and open observability across the fabric — all on open-standards hardware you choose. Swap suppliers, mix accelerators, adopt each hardware generation as it ships. The operating model stays the same.
The hyperscalers proved that AI needs a new network. They also proved you should never let one vendor own it.
1 min read
Every enterprise we talk to is somewhere on the same road. The faces in the boardroom change, the industry changes, the model du jour changes — but...
1 min read
When we look back at the history of networking, it's clear that the industry moves in distinct epochs. These technological eras aren't merely defined...
1 min read
Tom Hollingsworth the Networking Nerd asks, "Is Hedgehog the network OS distro?"