The Challenge

FarmGPU is a neocloud provider that builds and operates sustainable GPU cloud infrastructure and data centers for artificial intelligence and machine learning workloads.

FarmGPU set out to build its first B200 GPU cluster, expecting a straightforward six-server deployment that could be racked and cabled in an afternoon. Instead, the team hit a wall: proprietary InfiniBand pricing was roughly 3x the cost of an open Ethernet alternative, making it economically unviable for a neocloud competing on hourly GPU pricing. Choosing the cheaper open networking path introduced its own risk, though — nearly 90% of the total build time ended up going into network debugging rather than deployment. Optics wouldn’t initialize correctly, BIOS settings silently cut performance in half, NVIDIA driver bugs blocked CUDA on Blackwell, and inexperienced technicians miscabled ports. Every day of delay meant expensive B200s sitting idle instead of generating revenue for a company whose investors expected fast time-to-market.

The Solution

Hedgehog’s open-source network software and Celestica’s OCP-certified switches gave FarmGPU the access and partnership needed to debug, fix, and automate its solution.

FarmGPU partnered with Hedgehog for open-source AI network software and Celestica for OCP-certified DS5000 switch hardware, giving its engineers direct access to fix problems instead of waiting on vendor support tickets. When a critical optics initialization bug surfaced, engineers from Hedgehog, Celestica, and Broadcom joined multi-day live debugging sessions to trace the root cause together, something a closed proprietary stack would never have allowed. The team systematically resolved BIOS misconfigurations, NIC mode conflicts, and driver issues, then codified every fix into Ansible playbooks for repeatable future deployments. The result was a fully automated, validated 32-node B200 cluster running on an open Ethernet fabric at roughly half the cost of the InfiniBand alternative, launched commercially through RunPod’s Instant Clusters in 17 days.

 

 

At-a-Glance

FarmGPU_White_Full

 
COMPANY

FarmGPU


HEADQUARTERS

Rancho Cordova, CA


INDUSTRY

Neocloud

SCALE

32-node B200 cluster, 8x ConnectX-7 400G NICs per node, 400 GB/s bandwidth per node

HARDWARE

NVIDIA B200 GPUs, Celestica DS5000 (51Tb OCP switch), ConnectX-7 NICs, Supermicro servers

PRODUCTS USED

Hedgehog Open-source AI fabric software automating SONiC switch control plane for cluster and Celestica DS5000 switches running SONiC OS

WEBSITE

farmgpu.com

AI Networking Summit Video

  • FarmGPU discusses how they built a new AI cloud service powered by open networking solutions from Hedgehog and Celestica

  • FarmGPU showcases how open hardware and software from the Open Compute Project (OCP) ecosystem can deliver enterprise-grade AI infrastructure at cloud scale.

  • FarmGPU shares key test results comparing their OCP-based AI network against proprietary benchmarks, including Infiniband, across common AI workloads.

 

farmgpu transparent logo
Dig deeper by reading FarmGPUs blog post: Building an AI Cluster: Our 17-Day Crash Course in Open Networking 

Results

Achieved 50% Cost Reduction

Open Ethernet fabric cost roughly half of the proprietary InfiniBand alternative, freeing budget to buy more GPUs instead.

Eliminated Vendor Lock-in

Since Hedgehog runs on SONiC and works across white-box hardware, FarmGPU can mix vendors instead of being tied to a proprietary switch stack.

 

Elite Performance at Much Lower Cost

FarmGPU achieved 392 GB/s in NCCL all-reduce benchmarks — near-perfect scaling against a 400 GB/s theoretical max.

FarmGPU Performance graph


Source: SemiAnalysis

“The bottleneck has moved from compute to communication — whoever solves scale-up networking efficiently captures disproportionate value.”

 

 JM Hands, CEO and Co-Founder at FarmGPU

Additional Resources

 

 

In this AI Hedge episode, host Marc Austin (Hedgehog) talks with Jonmichael Hands (FarmGPU) about the rise of neoclouds—AI-first cloud providers built for training and inference. JM outlines why AI data centers need different infrastructure: dense power, liquid cooling, and fast networking/storage. They also discuss scaling challenges like reliability and supply chains, plus how OCP standards and ClusterMAX benchmarks lower deployment risk.