Interview

SemiAnalysis ClusterMAX 3.0: CoreWeave still the gold standard, Nebius reaches platinum, and why reliability gaps between neo clouds span minutes to days

Sep 24, 2026 with Jordan Nanos

Key Points

  • SemiAnalysis ranks CoreWeave the top neo cloud provider, but capacity constraints force $100M customers to wait until May 2025, opening an opening for Nebius, now elevated to platinum tier with feature-parity managed cluster quality.
  • Failure simulation testing reveals a stark reliability gap: top providers recover from induced hardware failures in 20 minutes using hot spare pools, while bottom-tier providers take hours or days and sometimes only respond after SemiAnalysis notifies them.
  • Inference platforms like Baseten and Modal are building captive chip capacity, while chip startups including Cerebras and SambaNova are becoming neo clouds themselves, forcing legacy infrastructure operators like Cloudflare and Akamai into the market.

Summary

SemiAnalysis ClusterMAX 3.0

SemiAnalysis has published ClusterMAX 3.0, its updated ranking of neo cloud providers, four to five months in the making. The focus this cycle is B300 GPUs and 800Gbps networks. CoreWeave holds the top position; Nebius has been elevated to platinum tier. SpaceX sits in an "unavailable" tier, pending testing.

CoreWeave vs. Nebius

CoreWeave retains the gold standard position on the strength of its monitoring infrastructure and operational scale — at peak, the company was deploying 10,000 GPUs a week. The problem is capacity: CoreWeave's balance sheet is full, and Jordan Nanos says customers with $100M to spend are being told they can't get allocation until May next year.

Nebius is capitalizing on that gap. Nanos argues that managed cluster quality from Nebius is now approaching feature parity with CoreWeave, with testing experience on both described as "really solid." The distinction is that CoreWeave and Nebius are both moving beyond managed clusters into bare metal deals and inference endpoints, where neither is the current leader.

“Top providers do this [failure recovery] in twenty minutes end to end, they have hot spare pools available. Some of the providers down the list — we simulate a failure or actually induce real hardware failures — and it'll sit there for hours, days, they don't identify it, we have to let them know. / CoreWeave has set the standard. They just have so much data for monitoring — deploying 10,000 GPUs a week at peak.”

Failure simulation testing

The most pointed addition to ClusterMAX 3.0 is hands-on failure simulation. SemiAnalysis induces real hardware failures and measures how long it takes each provider to identify and recover from them. The spread is stark: top providers resolve failures end-to-end in 20 minutes using hot spare pools. Providers at the bottom of the rankings can take hours or days, and sometimes only act after SemiAnalysis notifies them directly. That gap is the sharpest differentiator between providers that have roughly similar hardware on similar pricing.

Network design and software lag

Beyond reliability, network architecture is the next major performance variable. Providers using non-Nvidia network hardware — Amazon EFA is the named example — face a delay of months before new open-source inference software like vLLM or SGLang is properly supported on their custom stack. Providers on Nvidia's standard InfiniBand or RoCE avoid that lag, but the tradeoff is less flexibility in cluster design at scale.

Who builds a neo cloud next

Nanos sees three distinct incoming waves. Inference platforms like Baseten and Modal are moving toward owning chips outright, driven by customer demand. Chip startups are becoming neo clouds themselves — Cerebras, Groq's remaining entity after an Nvidia acquihire, SambaNova, and OpenAI's Jalapeno self-build program are all cited. And legacy infrastructure operators with no current GPU presence — Cloudflare, Akamai, Iron Mountain — are entering the space.

On talent, electrician and data center technician salaries in markets like Louisiana and Abilene, Texas are running three to five times prior rates. Crusoe is named as a driver of that compression. The neo clouds winning on staffing are cross-training software engineers into SRE roles and retraining ground-level technical workers rather than purely competing for existing talent.

Mixed GPU strategy

The viable path on multi-vendor GPU strategy is to partner exclusively with Nvidia on the way up, then add alternative suppliers once scale creates negotiating leverage. Trying to run a mixed-chip fleet from the start is, in Nanos's view, a losing position — Nvidia's allocation preferences can make or break a business, and you need that partnership to grow.

The workload shapes the hardware requirement more than most operators acknowledge. Pre-training clusters, inference clusters, and RL infrastructure each have materially different optimal designs. Nanos flags CPU and non-GPU compute as underrated for RL training, where data centers full of chips simulating phone and laptop behavior during training are becoming a real infrastructure category.

SemiAnalysis says it will be testing inference endpoints, RL infrastructure, and alternative chip providers in the coming months.

Every deal, every interview. 5 minutes.

TBPN Digest delivers summaries of the latest fundraises, interviews and tech news from TBPN, every weekday.