Case Study
FarmGPU achieved near-perfect performance at half the cost with Hedgehog
The Challenge
FarmGPU is a neocloud provider that builds and operates sustainable GPU cloud infrastructure and data centers for artificial intelligence and machine learning workloads.
FarmGPU set out to build its first B200 GPU cluster, expecting a straightforward six-server deployment that could be racked and cabled in an afternoon. Instead, the team hit a wall: proprietary InfiniBand pricing was roughly 3x the cost of an open Ethernet alternative, making it economically unviable for a neocloud competing on hourly GPU pricing. Choosing the cheaper open networking path introduced its own risk, though — nearly 90% of the total build time ended up going into network debugging rather than deployment. Optics wouldn’t initialize correctly, BIOS settings silently cut performance in half, NVIDIA driver bugs blocked CUDA on Blackwell, and inexperienced technicians miscabled ports. Every day of delay meant expensive B200s sitting idle instead of generating revenue for a company whose investors expected fast time-to-market.
The Solution
Hedgehog’s open-source network software and Celestica’s OCP-certified switches gave FarmGPU the access and partnership needed to debug, fix, and automate its solution.
FarmGPU partnered with Hedgehog for open-source AI network software and Celestica for OCP-certified DS5000 switch hardware, giving its engineers direct access to fix problems instead of waiting on vendor support tickets. When a critical optics initialization bug surfaced, engineers from Hedgehog, Celestica, and Broadcom joined multi-day live debugging sessions to trace the root cause together, something a closed proprietary stack would never have allowed. The team systematically resolved BIOS misconfigurations, NIC mode conflicts, and driver issues, then codified every fix into Ansible playbooks for repeatable future deployments. The result was a fully automated, validated 32-node B200 cluster running on an open Ethernet fabric at roughly half the cost of the InfiniBand alternative, launched commercially through RunPod’s Instant Clusters in 17 days.
At-a-Glance
COMPANY
FarmGPU
HEADQUARTERS
Rancho Cordova, CA
INDUSTRY
Neocloud
SCALE
32-node B200 cluster, 8x ConnectX-7 400G NICs per node, 400 GB/s bandwidth per node
HARDWARE
NVIDIA B200 GPUs, Celestica DS5000 (51Tb OCP switch), ConnectX-7 NICs, Supermicro servers
PRODUCTS USED
Hedgehog Open-source AI fabric software automating SONiC switch control plane for cluster and Celestica DS5000 switches running SONiC OS
WEBSITE
AI Networking Summit Video
-
FarmGPU discusses how they built a new AI cloud service powered by open networking solutions from Hedgehog and Celestica
-
FarmGPU showcases how open hardware and software from the Open Compute Project (OCP) ecosystem can deliver enterprise-grade AI infrastructure at cloud scale.
-
FarmGPU shares key test results comparing their OCP-based AI network against proprietary benchmarks, including Infiniband, across common AI workloads.
Results
Achieved 50% Cost Reduction
Open Ethernet fabric cost roughly half of the proprietary InfiniBand alternative, freeing budget to buy more GPUs instead.
Eliminated Vendor Lock-in
Since Hedgehog runs on SONiC and works across white-box hardware, FarmGPU can mix vendors instead of being tied to a proprietary switch stack.
Elite Performance at Much Lower Cost
FarmGPU achieved 392 GB/s in NCCL all-reduce benchmarks — near-perfect scaling against a 400 GB/s theoretical max.
Source: SemiAnalysis
“The bottleneck has moved from compute to communication — whoever solves scale-up networking efficiently captures disproportionate value.”
JM Hands, CEO and Co-Founder at FarmGPU
Additional Resources
In this AI Hedge episode, host Marc Austin (Hedgehog) talks with Jonmichael Hands (FarmGPU) about the rise of neoclouds—AI-first cloud providers built for training and inference. JM outlines why AI data centers need different infrastructure: dense power, liquid cooling, and fast networking/storage. They also discuss scaling challenges like reliability and supply chains, plus how OCP standards and ClusterMAX benchmarks lower deployment risk.