What Neocloud Providers Actually Need from Infrastructure Vendors

The AI infrastructure conversation has become obsessed with GPUs. 

Every major announcement seems to revolve around who’s secured the latest accelerator, who’s deploying the largest cluster, or who raised the most money to build the next AI data center. GPUs have become the headline, but they’re only part of the story.

The industry’s challenge is turning those GPUs into reliable, scalable, and profitable AI services, and this is the challenge neocloud providers are solving.

Neocloud providers aren’t just GPU capacity brokers renting out compute. They represent a new category of AI infrastructure companies, including AI-native public clouds, sovereign AI infrastructure builders, private AI cloud operators, GPU marketplaces, and inference platforms. They differ in maturity and business model, but they share one reality: their success depends on turning scarce, capital-intensive GPU and AI infrastructure into reliable, differentiated, highly-utilized services.

And this is what they need from infrastructure vendors.

Whereas traditional cloud infrastructure is evaluated around general-purpose scale, price, and availability, neocloud infrastructure is different. These providers are building dense GPU environments where power, cooling, networking, optics, storage, orchestration, and operational tooling all affect customer experience and financial margins. Companies like CoreWeave and Crusoe have demonstrated that building AI infrastructure is not simply about assembling servers. It requires coordinating dense GPU environments, advanced cooling, high-performance networking, energy strategy, and operational expertise into a single platform. 

For infrastructure vendors who want to service these kinds of companies well, the message is clear that neoclouds do not want isolated products. They want infrastructure building blocks that help them scale faster, operate better, preserve flexibility, and improve their bottom line.

AI Workloads Aren’t All the Same

One of the biggest mistakes infrastructure vendors can make is treating AI as a single workload.

The phrase “AI workload” is too broad to be useful. Large-scale model training, fine-tuning, inference, retrieval-augmented generation, batch experimentation, and enterprise AI deployments each stress infrastructure in different ways.

Training LLMs can be dominated by collective communication and east-west GPU traffic. Inference can be latency-sensitive, geographically distributed, and increasingly tied to storage and cache movement. And adding another layer of complexity, fine-tuning and enterprise private AI often require stronger isolation, support models, and predictable performance.

Infrastructure vendors that succeed in this market are the ones that understand how different workloads place different demands on networking, automation, and observability. One example is Lambda, whose infrastructure messaging highlights bare-metal instances, custom networking, and system-level optimizations for distributed training workloads. 

Infrastructure vendors need to stop assuming one reference architecture solves every neocloud problem. This new breed of cloud providers need infrastructure designed around their actual workload mix, with appropriate fabric design, optics, storage networking, telemetry, and automation.

Networking for GPU Economics

In neocloud environments, the network is less like plumbing and more of a utilization engine. When GPUs cost tens of thousands of dollars each and clusters of GPUs scale into the thousands, every percentage point of lost utilization has a direct financial impact. 

This is why AI networking has become a battleground. Technologies like NVIDIA's Spectrum-Xplatform and the work being done by the Ultra Ethernet Consortium reflect an industry-wide effort to optimize Ethernet for the communication patterns generated by large-scale AI workloads. 

Neocloud operators need infrastructure vendors to prove performance across real conditions. 

  • How does the environment behave during large AllReduce operations? 

  • What happens when multiple tenants compete for resources?

  • How does the fabric handle checkpointing, storage-intensive inference, or partial cluster expansion?

  • How quickly can operators identify congestion before customers notice performance degradation?

These are operational questions, not marketing questions, and the real question is “Does the fabric keep GPUs productively occupied across the range of services this neocloud sells?”

Open Architectures Need Operational Maturity

Many neoclouds want alternatives to vertically integrated stacks, but they want to avoid becoming their own unsupported systems integrator. Open Ethernet, SONiC, UEC alignment, merchant silicon, and modular architectures can preserve vendor leverage and future flexibility. 

But openness without strong validation, support, observability, and lifecycle management increases operational risk. NVIDIA-aligned architectures can reduce integration friction, but they also raise questions about supplier concentration and service differentiation.

This is where infrastructure vendors have an opportunity. The strongest message isn’t “open networking” as an ideology. It’s “choice with accountability”, giving customers architectural flexibility without forcing them to become full-time systems integrators. 

Neoclouds need vendors that can support flexible architectures, heterogeneous accelerator futures, multiple operating models, and mixed topologies while still providing operational reliability as clusters grow and evolve.

Power is Part of the Architecture

For many AI providers, power has become the limiting resource - not compute.

Companies like Nebius have announced AI campuses measured in hundreds of megawatts, while providers like Crusoe have built their market position around access to affordable energy and purpose-built infrastructure. These announcements highlight an important trend in today’s industry: power strategy is rapidly becoming infrastructure strategy.

This also means infrastructure decisions can no longer happen in isolation. Network architecture, rack design, cooling strategy, and power delivery are becoming tightly connected. A networking decision that reduces power consumption at scale can influence facility design. A higher-density rack design can change cooling requirements. The entire infrastructure stack is becoming part of the energy strategy. 

As clusters extend across rooms, buildings, campuses, and metro areas, optics and data center interconnect also become strategic. Vendors that can connect backend fabric, routing, optical transport, telemetry, and automation into a coherent scale-out and scale-across story will be more valuable than vendors selling single-domain products.

Building Both Faster and Better

Neocloud providers operate under intense timing pressure. Large customer contracts, GPU delivery windows, power availability, and financing all create urgency. If capacity is delayed, revenue is delayed. At the same time, bringing capacity online before operational readiness creates its own financial risks.

That means infrastructure vendors need to help reduce deployment uncertainty. Validated designs, realistic lead times, predictable supply, pre-tested optics, automation templates, and clear sparing models matter. So do commercial models that reflect bursty growth rather than slow enterprise refresh cycles.

The best vendors will help neoclouds answer a much more important question: “How predictable will this vendor be and how much do they understand our timing issues?”

Differentiation Over Scale

Not every neocloud is trying to become the next hyperscaler. If every vendor uses the same architecture, sells the same GPUs, and competes only on price, margins will compress. Providers need room to differentiate by workload specialization, regional presence, sovereign posture, private-cluster capabilities, performance consistency, energy strategy, security, and operational experience.

A provider specializing in distributed model training will likely prioritize high-performance backend fabrics and collective communication efficiency, whereas an enterprise-focused provider emphasizes workload isolation, observability, compliance, and operational consistency. 

The infrastructure underneath these services has to enable differentiation, not eliminate it. A vendor that forces every customer into the same architecture may simplify deployment, but it also limits the provider's ability to create unique offerings. 

Rather than leading with predefined reference architectures, infrastructure vendors should begin by understanding the business their customer is trying to build. The most valuable conversations aren’t about individual products, but instead about helping providers create services that are faster, more reliable, more efficient, or more differentiated than their competitors. 

Infrastructure Vendors Need to Think Differently

The neocloud market is still in its early stages, which means the rules are still being written. For infrastructure vendors, winning in the neocloud market requires a new mindset. The strongest partners will bring four things to the table: 

First, they need technical specificity. They need workload and business model-specific designs, not generic AI claims.

Second, they need operational capabilities. This includes telemetry, automation, congestion visibility, and support for day-two realities.

Third, they need architectural flexibility, or in other words, support for open standards, heterogeneous futures, and growth that does not always happen in perfect reference-architecture increments.

And lastly, they need business empathy, which is an understanding that power, capital, utilization, customer concentration, and time-to-revenue are as important as feeds and speeds.

Solutional’s working thesis is that the neocloud market is entering a sorting phase. The winners will not simply be the providers with access to GPUs - they’ll be the operators that can provide GPUs and infrastructure as reliable, efficient, differentiated services. 

The next phase of AI infrastructure will not only be defined by who can acquire the most GPUs. It will be defined by who can turn those GPUs into dependable, scalable, and differentiated services.

Scott Robohn

Scott is co-founder and CEO of Solutional, where he leads initiatives in next-gen networking, automation, AI, and emerging technologies. With 35+ years of experience building, guiding, and scaling technical teams and solutions, Scott helps IT and NetOps organizations evolve into software-centric, resilient, and intelligent operations teams. His career has provided the depth and breadth of experience needed to lead technical sales organizations, including roles and partnerships with CTOs, Solutions Architects, Sales Engineers, Systems Engineers, Account Executives, Product and Engineering leaders, and other job functions. Scott is a frequent event speaker, host of the Total Network Operations podcast, and a co-founder of the Network Automation Forum (NAF).

Next
Next

Not Every Network Problem Needs an LLM