Why GPU Clouds Are Not Just Smaller Public Clouds

The fastest way to misunderstand neocloud providers is to describe them as smaller versions of AWS, Microsoft Azure, or Google Cloud.

That framing is tempting. Neoclouds sell cloud infrastructure, and many offer GPU instances, storage, networking, Kubernetes, managed services, and enterprise support. Some even have hyperscalers as customers or partners. But the underlying business is different enough that vendors, investors, and customers need to treat the neocloud category on its own terms.

What is a Neocloud?

A neocloud provider provides a GPU infrastructure, often consumed like a public cloud. Neclouds are making a much more concentrated bet: that specialized AI infrastructure can be built, operated, and sold better as a dedicated infrastructure and service than it can be through a general-purpose cloud. 

Public cloud providers and many hyperscalers are usually built and operated as a broad general-purpose cloud platform that monetizes thousands of services across compute, storage, databases, networking, security, analytics, AI, collaboration, and enterprise software ecosystems. There is very little specialization as is the case with neoclouds, and that difference matters.

Neoclouds begin with scarcity, not breadth

The hyperscaler model is built on breadth, global reach, and platform gravity. The neocloud model emerged from a more specific market scenario. High-end GPU capacity was scarce, AI demand was exploding, and many customers needed faster and more specialized access than traditional cloud channels could provide.

McKinsey describes neoclouds as independent GPU-as-a-service providers that emerged in response to global scarcity of high-end compute and the revenue-diversification strategies of major chip producers. They also note that the original bare-metal GPU-as-a-service economics can be fragile. That is an important distinction because neoclouds were not created to replicate the full hyperscaler service catalog. Instead, they were created to solve acute AI infrastructure constraints (McKinsey & Company).

This is why many neoclouds sound less like general-purpose cloud companies and more like infrastructure specialists. For example, CoreWeave positions itself as an AI-native cloud platform purpose-built for complex AI workloads. Lambda emphasizes dedicated bare-metal GPU clusters, low-latency networking, and high-throughput interconnects for distributed AI workloads. RunPod highlights on-demand GPUs and serverless compute for training, inference, and batch AI workloads (CoreWeave).

The center of gravity is not “cloud services.” It’s useful, available, and scalable GPU capacity.

Neocloud infrastructure requirements are more specialized

Public clouds optimize for many workloads at once. GPU clouds optimize around a narrower but extremely demanding set of workloads: model training, fine-tuning, inference, RAG pipelines, synthetic data generation, batch experimentation, and private AI clusters. That drives different infrastructure requirements.

A general-purpose cloud region may be judged by service breadth, availability zones, storage durability, developer ecosystem, enterprise contracts, and compliance coverage. A neocloud cluster focuses on GPU availability, job start time, accelerator generation, interconnect performance, storage throughput, orchestration model, workload isolation, and the provider’s ability to keep expensive accelerators highly utilized.

Nebius, for example, markets its AI cloud around current NVIDIA GPU generations, InfiniBand networking, managed Kubernetes, Slurm-based clusters, and fast storage. That’s a very different public message from a general-purpose cloud provider leading with a broad catalog of enterprise services (Nebius).

This specialization creates opportunity, but it also creates risk. A GPU cloud’s infrastructure decisions are tightly coupled to its economics. If the network slows training jobs, if storage can’t keep up, if power delivery is delayed, or if capacity sits idle, the business impact is immediate.

Neoclouds are more exposed to capital timing

Public clouds are capital-intensive, but they usually have diversified revenue streams, large balance sheets, and deep customer ecosystems. Neoclouds may have significant demand, but they often operate with more concentrated exposure to GPU supply, debt, lease commitments, customer concentration, and rapid infrastructure depreciation.

CoreWeave is the obvious public example. Recent reporting describes extraordinary growth and a large revenue backlog, but also very high capex expectations, debt load, lease liabilities, depreciation, interest expense, and sensitivity to component costs (The Wall Street Journal).

This doesn’t mean the model is flawed. It means the business model is unforgiving. A neocloud has to turn capital into deployed capacity, deployed capacity into contracted workloads, and contracted workloads into profitable utilization quickly. That makes infrastructure vendors more than suppliers and instead a part of the provider’s time-to-revenue equation.

Power and site strategy are first-order business issues

For most public clouds, power is a major constraint. For neoclouds, power can actually be the strategy. Crusoe is a great example. It describes its AI data centers as purpose-built for high-performance workloads, with advanced cooling, networking, and an energy-first approach. McKinsey’s interview with Crusoe’s CEO describes the company’s focus on building where low-cost, abundant, or stranded energy resources are available (Crusoe).

Nscale is another good example also positioning around power, platform, and scale, reporting that it’s combining data centers, GPU fleets, and energy development as part of its AI factory strategy (Data Center Frontier).

This is another reason neoclouds are not merely smaller public clouds. Their differentiation may actually come from power access, site selection, cooling design, modular deployment, or sovereign/regional infrastructure strategy as much as from software features.

They must differentiate before capacity commoditizes

GPU scarcity created the first wave of neocloud opportunity, but scarcity alone is not a durable moat. As more capacity enters the market, customers will ask harder questions:

  • Which provider delivers the best job completion time? 

  • Which one supports the right accelerator mix? 

  • Which one offers reliable private clusters? 

  • Which one has predictable inference economics? 

  • Which one can support sovereign requirements? 

  • Which one has better networking, storage, observability, and support?

This is where neoclouds diverge from public clouds yet again. Public cloud providers can retain customers through broad platform integration, but neoclouds need to earn loyalty through workload-specific infrastructure performance, commercial flexibility, and operational trust.

Vultr’s public analysis of neocloud consolidation argues that the market is entering a period where capital, scale, and enterprise capabilities will shape which GPU providers survive (Vultr Blogs). Even if one disagrees with the timing or specific winners, the underlying point is sound: access to GPUs is becoming necessary but insufficient.

Their vendor needs are different

A cloud provider may design, build, and operate much of its infrastructure stack internally. A smaller GPU cloud may need infrastructure vendors to provide more complete building blocks, but without locking the provider into a rigid model that prevents differentiation. That creates a specific vendor requirement: validated flexibility.

Neoclouds need proven designs, reliable supply, automation, telemetry, congestion management, optics, routing, and support. But they also need freedom to support different customer types, workload patterns, accelerator generations, operating models, and geographic expansion plans.

The better mental model

The better model is this:

  • A public cloud provider is a platform economy.

  • A neocloud is an AI infrastructure operating company.

The difference between the two is really based on what matters to each organization. 

This new type of provider, the neocloud, needs power, GPUs, networking, optics, storage, orchestration, observability, and customer demand to line up in a narrow window. Neoclouds need to scale quickly without overcommitting to architectures that limit future flexibility. They need to sell differentiated services before raw GPU access becomes more commoditized. And they need infrastructure vendors that understand the business consequences of technical decisions.

Neoclouds, and GPU clouds more broadly, may eventually look more like public cloud providers in some areas. They may add more managed services, developer tooling, compliance capabilities, and enterprise support. But their starting point is different, their risk profile is different, and their infrastructure priorities are different.

Vendors that understand this will have better conversations. Vendors that don’t will show up trying to sell a smaller version of a public cloud solution into a market that is being shaped by very different constraints. 

Scott Robohn

Scott is co-founder and CEO of Solutional, where he leads initiatives in next-gen networking, automation, AI, and emerging technologies. With 35+ years of experience building, guiding, and scaling technical teams and solutions, Scott helps IT and NetOps organizations evolve into software-centric, resilient, and intelligent operations teams. His career has provided the depth and breadth of experience needed to lead technical sales organizations, including roles and partnerships with CTOs, Solutions Architects, Sales Engineers, Systems Engineers, Account Executives, Product and Engineering leaders, and other job functions. Scott is a frequent event speaker, host of the Total Network Operations podcast, and a co-founder of the Network Automation Forum (NAF).

Next
Next

NVIDIA, Hugging Face, and the Growing Importance of Open-Weight AI