Hedgehog AI Network

Podcast - FarmGPU on the Ticking Clock Driving Neocloud Economics

Written by Marc Austin | Jul 21, 2026 10:57:03 PM

A megawatt GPU cluster can cost roughly $1 million to set up, and every month that build runs late is about $1.4 million in lost revenue. That unforgiving math sits at the center of my latest conversation on The AI Hedge with Jonmichael Hands, founder and CEO of FarmGPU. Jonmichael and I have shared a stage before to talk about networking, but this time I wanted his read on neo-clouds and what it takes to operate AI infrastructure at scale.

What Is a Neocloud?

A neo-cloud, Jonmichael explained, was shorthand for a cloud provider built specifically to serve AI workloads, something traditional hyperscalers weren’t originally designed to do at scale. He pushed back on the idea that neo-clouds were cheaper clones of AWS, Azure, and Google Cloud, with different neo-clouds driving different platform requirements. Providers like RunPod, one of FarmGPU’s partners, have built a business around 500,000 developers and curated a really nice user experience for developers to spin up a pod really fast within a few seconds.

Others focus on renting massive clusters to frontier labs for training and production scale inference. Treating neo-clouds as interchangeable is inaccurate.

Why AI Data Centers Aren’t Traditional Data Centers

Jonmichael pointed to three things that separate AI infrastructure from traditional data centers: power density, liquid cooling, and high-speed networking. AI data centers can draw power of 100 to 140 kilowatts per rack, making liquid cooling a must.

Training clusters also split traffic across a back-end fabric connecting GPU to GPU and a front-end fabric handling storage and internet traffic, with NVLink moving data at roughly 1.8 terabytes per second inside a rack. As inference scales in production, storage bandwidth matters just as much.

Grading the Market With ClusterMax

We spent a good chunk of the conversation on ClusterMax, the rating system from SemiAnalysis that scores neo-clouds across security, lifecycle and orchestration, storage, reliability, and networking.

According to Jonmichael, storage is where FarmGPU differentiates, unsurprising given his background running Intel’s Data Center SSD product line and contributing to the NVMe and OCP SSD specifications.

On reliability, the company built two internal tools: Haystack, an observability platform pulling switch telemetry from Hedgehog’s open API into Prometheus dashboards, and Shepherd, an autonomous SRE that reads system and kernel logs to catch failures before a customer notices. GPU servers fail in multiple ways, NVLink and switch failures to driver crashes, which is why Jonmichael’s leaning hard on automation and AI agents to keep up.

The Clock That Never Stops

The part of the conversation that stuck with me most was the economics of time-to-revenue. Leasing space runs around $150 per kilowatt per month, meaning a megawatt cluster can carry a $150,000 monthly shell lease before a single watt of power gets billed to a customer.

“The dominating part of TCO and OpEx is actually the data center shell lease,” Jonmichael said. Fall three months behind, and a neo-cloud operator is looking at roughly $1.4 million in lost monthly revenue, on top of $450,000 in operating costs and a rapidly depreciating asset sitting idle.

Betting on Open Standards

FarmGPU was the first customer for a new OCP reference architecture for training and inference networking, built with Hedgehog, and Jonmichael now co-chairs an OCP working group on scaling AI clusters for neo-clouds.

“My favorite thing about OCP is when they say open, it really is open. Every single meeting is recorded and posted on YouTube. All the meeting notes are available in a shared Google doc,” he said. For smaller operators, validated reference designs cut deployment risk and reduce vendor lock-in.

Key Takeaway

Neoclouds are not a monolith, and treating them as commodity GPU rentals misses the whole story. Winning depends on solving problems hyperscalers have never faced at this scale: extreme power density, liquid cooling, storage built for inference, and reliability engineering for hardware that fails often and unpredictably.

Jonmichael’s experience shows a small team can compete on operations and automation, and that open standards like OCP’s reference architectures are becoming the shortcut smaller players need to reduce risk. In a market racing against the clock, speed, openness, and operational discipline will define the neo-clouds built to last.

Listen to the full episode here.