
That is the boundary. Hyperscalers still fit when identity, data gravity, compliance mappings, and dozens of managed services must sit in the same account. A neocloud is also distinct from colocation with a GPU invoice attached. Billing, tenancy, and a control plane are cloud-shaped. The frequent mix-up is treating every H100 rental as the same product. Marketplace GPU hosts, hyperscaler GPU instances, and neocloud clusters share silicon. They differ in fabric design, reservation terms, capacity guarantees, and how much of the stack the customer has to run.
That distinction matters because a lower GPU-hour price does not necessarily produce a lower project cost. Data movement, orchestration, networking, reserved-capacity commitments, and the engineering required to operate the environment can all change the economics.
Over the next 12 to 24 months the category will be tested less by new listings and more by whether bursty, latency-sensitive inference can pay for the dense fabrics that training reserved. Hyperscalers will keep pulling those jobs back with custom accelerators and tighter coupling to their data platforms. Power and NVIDIA allocation remain the hard limits. Software catalogs do not.
A Specialist Cloud, Not a Smaller AWS
Count is not enough. You need a single GPU generation, a documented rail design, and a fabric that keeps collective operations inside the reservation. Multi-node training is limited by interconnect. Ask for topology, oversubscription, and whether the block is contiguous or a best-effort pool. Look for measured all-reduce times at the size you will run.
Nutanix describes a hybrid routing habit that matches how many IT shops already buy: send the dense AI job to the platform that can run it, keep the rest of the estate where it already lives.
A 64-GPU run on a neocloud is typically a job against a reservation ID. The same run on a hyperscaler often waits on instance quota and a VM family that has to live in the same region as the rest of the estate. Power, cooling, and the provider’s NVIDIA allocation sit underneath both stories. Those constraints decide whether the cluster is real.
Clusters, Fabrics, and the Job That Fills Them
What operational work comes back in-house?
The market is splitting by how providers obtain GPUs, power, and customers.
CoreWeave is the public reference for large reserved training clusters after its 2025 IPO; ABI Research notes cluster work supporting OpenAI and Microsoft. Lambda competes as an AI cloud with aggressive on-demand GPU pricing, including the B200 rate above. Crusoe sites capacity near energy sources and, unlike several NVIDIA-only peers, lists AMD MI300X and MI355X on its rate card. Nebius, Amsterdam-headquartered and Nasdaq-listed, has used European and sovereign-cloud positioning, including a Microsoft partnership described by ABI Research. CNBC reported Nebius revenue of 5 million in that August 2026 cycle, up 514%.
Synergy Research Group put 2025 neocloud revenue above billion, with fourth-quarter revenue at billion, up 223% year over year. CoreWeave’s March 2025 Nasdaq listing under ticker CRWV pulled the category onto finance-team agendas, not only ML ops runbooks. CIO Dive reported that CoreWeave, Vultr and FluidStack, are still chasing enterprise AI infrastructure demand into 2026.
Training and large fine-tunes reserve a block of identical GPUs, often eight per node, joined by InfiniBand or another RDMA fabric so gradient exchange does not stall on ordinary Ethernet. The provider images the nodes and attaches fast storage. Inputs are the model code, a dataset sitting in object storage, and a reservation size and duration. The customer runs PyTorch, JAX, or a similar framework, with Slurm or Kubernetes as the scheduler. Outputs are checkpoints, metrics, and logs. The control point is topology. If all-reduce traffic leaves the intended rail, utilization falls even while the GPUs show as allocated.
Public Listings, Inference, and Hyperscaler Pushback
CoreWeave now reports as a public company. In August 2026 it posted second-quarter revenue of .6 billion, up 112% from the year-earlier quarter. Those are operating results, not launch-day bookings.
The claim that neoclouds take over general IT remains speculative. Current buying still parks systems of record on AWS, Azure, or Google. The contest that is actually funded is over dense GPU hours, reservation length, and whether hyperscalers close the gap with custom chips and their own clusters.
Enterprise teams that first used neoclouds as overflow for training are now placing production inference on some of the same providers. That shift is emerging, not fully settled. In April 2026 Nutanix said it would add multitenant capabilities aimed at neoclouds that deliver AI services. Energy-sited and sovereign capacity, already part of how Crusoe and Nebius sell, is likely to weigh more over the next 12 to 24 months as residency rules tighten.
Who holds the NVIDIA allocation if demand tightens?
Three Questions Before You Sign a GPU Reservation
More than on a hyperscaler GPU instance if you take bare metal. Identity, patching, orchestration, observability, and data egress become yours unless the provider sells a managed layer. Bare metal is often how the performance is delivered. Evaluate IAM federation, private connectivity to your existing cloud, encryption, and audit logs against your control framework. If the team cannot operate a GPU cluster, a lower hourly rate does not cut cost.
Inference uses the same chips in smaller slices, sometimes mixed with CPUs, to serve tokens against latency targets. Batching, autoscaling, and model placement matter more than finishing a 3,000-GPU job. Providers that sold reserved training blocks have been adding on-demand inference because idle reserved GPUs waste capital. Isolation is harder than in a classic VM cloud: GPU and RDMA tenancy is less mature, so noisy neighbors and firmware discipline become part of the purchase, not an afterthought.
Your reservation contract and the provider’s supply. Intent to burst later does not create silicon. Neoclouds sit downstream of a concentrated accelerator vendor. If their allocation slips, your queue returns. Look for committed capacity language, substitution rights across GPU generations, and expiry terms. Published on-demand rates help with overflow. In a 2026 GPU-cloud comparison, Lambda listed on-demand B200 at .69 per GPU-hour. That is a rate card. It does not replace a multi-month training block.
When a training job needs hundreds of tightly coupled GPUs, the rest of a cloud catalog is noise. Extra regions and managed databases do not help if accelerator inventory is sold out. Neoclouds exist for that condition. They rent GPU clusters as the product, not as one SKU among hundreds.
Confirmation would look like multi-year enterprise reservations that are more than overflow from a single foundation-model lab, inference SLOs that hold under production traffic, and customers who keep training on a neocloud while leaving identity and systems of record on AWS, Azure, or Google. Challenge would look like sustained large-cluster GPU availability at hyperscalers on comparable terms, or buyers returning after fabric and capacity misses.
Will the cluster topology match the job, or only the GPU count?
Capacity Brokers, Energy Sites, and Sovereign Bets
Two paths show how the product actually behaves.
For enterprises, that makes portability increasingly important. A model may be trained on one provider, moved to another for inference, and connected back to data and identity services that remain in the primary cloud. The technical question is no longer simply where the GPU is cheapest. It is whether the workload can move without creating a new operational dependency.
A familiar pattern follows. The company already has AWS, Azure, or Google Cloud. Identity, data, and most applications stay there. The next pretraining or fine-tune still cannot get a dense block of H100 or B200 GPUs on the schedule the model team needs. The live constraint is queue time and cluster shape.
Watch Inference Mix, Not Another IPO
A neocloud is a cloud provider organized around GPU-as-a-service for AI training, fine-tuning, and inference. The usual offer is bare-metal or thin-VM access to current NVIDIA accelerators, and in some cases AMD, plus high-bandwidth GPU-to-GPU networking and storage that can feed the job. Network World describes CoreWeave, Lambda Labs, Crusoe, and Nebius as competing with AWS, Azure, and Google on GPU processing-as-a-service. The hyperscalers still sell the general-purpose platform.
The live decision inside most organizations is smaller than the market story. It is whether the next training run or inference fleet is limited by catalog comfort or by GPUs that actually show up.
That supply-side model is increasingly important. The neocloud is not simply buying GPUs and reselling compute. It is assembling power, cooling, networking, accelerator inventory, data-center capacity, and financing into a service that can be sold against a specific workload. The result is a market where access to physical infrastructure can be as important as software features.






