There Is No Standard AI Factory

Spend enough time around AI infrastructure and you start hearing people talk about “the AI factory” as if everyone is building roughly the same thing. The reality is that’s partly true, but not entirely.

There are certainly common patterns. High-density GPU systems. High-bandwidth scale-up connectivity. High-performance scale-out networks. Fast storage. Separate management networks. Sophisticated orchestration and telemetry.

Reference architectures have also become enormously valuable. NVIDIA, in particular, has put substantial engineering into reference architectures that give operators tested designs for building GPU infrastructure without starting from a blank sheet of paper. In fact, to be an NVIDIA Cloud Partner requires adhering to NVIDIA’s architectural standards. 

But a reference architecture is not evidence that every AI factory should look the same.

The specific architecture depends on what GPUs are being deployed, which generation of those GPUs is involved, what workloads the infrastructure will run, how large the environment needs to become, how customers consume it, and how the operator expects the business to evolve.

That makes designing an AI factory less like selecting a standard data center architecture and more like designing a system around a particular set of technical and business requirements.

Even NVIDIA AI Factories Are Changing

One reason it is dangerous to think of AI infrastructure as a fixed architecture is how quickly the underlying systems are changing.

For example, consider NVIDIA GB300 NVL72.

A GB300 NVL72 rack contains 72 Blackwell Ultra GPUs and 36 Grace CPUs. Within the rack, fifth-generation NVLink provides the high-bandwidth scale-up domain. Outside that domain, NVIDIA’s current enterprise reference architecture uses ConnectX-8 SuperNICs for the east-west compute fabric, with Spectrum-X Ethernet as one supported networking platform. NVIDIA also supports Quantum-X800 InfiniBand for GB300 deployments.

The resulting network is not simply “a faster leaf-spine.” The rack itself has become a tightly integrated compute system, and the external network extends that system across racks.

However, here comes Vera Rubin.

The networking architecture changes with Vera Rubin GPUs. NVIDIA’s Rubin platform introduces sixth-generation NVLink for scale-up communication and ConnectX-9 SuperNICs with up to 1.6 Tbps of scale-out network bandwidth per GPU. Spectrum-6 and the latest Spectrum-X architecture are designed alongside those systems.

This means that a network designed around one generation of GPU infrastructure can’t automatically be assumed to be the right design for the next. NIC bandwidth changes, as do scale-up domains, switch generations, optical requirements, power and cooling constraints, and the relationship between scale-up and scale-out.

In effect, the GPU roadmap is also a network roadmap.

NVIDIA Is Not the Only Architecture

There’s another problem with treating “AI factory” as shorthand for one architecture: not every AI factory is built around NVIDIA GPUs.

AMD’s Helios rack-scale architecture is a good example. Helios combines 72 AMD Instinct MI455X GPUs with AMD EPYC CPUs, Pensando networking, ROCm software, and an architecture built around open industry standards. AMD is also developing its networking stack across multiple layers, including scale-up connectivity, scale-out networking, and front-end infrastructure.

That produces a different set of architectural decisions than an NVIDIA NVL72 environment, and this is likely to become more important rather than less.

Operators may choose different accelerators because of availability, economics, customer demand, workload characteristics, software ecosystem, power efficiency, or a desire to avoid dependence on a single supplier. Large operators may eventually have several generations and types of accelerators operating at the same time, and the network architecture has to accommodate that reality.

The question is therefore not simply, “How do we build an AI network?” Instead, it’s, “What kind of AI infrastructure are we actually building?”

Training Changes the Network

Workload is another major variable. Large-scale distributed training is probably the workload most people picture when they think about AI networking.

Thousands of GPUs may participate in the same job. Collective operations continually exchange data between accelerators and the system repeatedly computes, communicates, and synchronizes. At that point, the scale-out network effectively becomes part of the compute system.

Congestion that delays GPU communication can delay synchronization, and delayed synchronization can leave expensive GPUs sitting idle. That can extend job completion time and reduce the amount of useful work the infrastructure produces.

For a training-heavy AI factory, this puts a huge emphasis on predictable bandwidth, low latency, congestion management, topology, RDMA performance, failure recovery, and visibility into the relationship between network behavior and application performance.

In other words, the operator needs to understand whether the network is allowing the GPU cluster to operate efficiently as a distributed system.

Inference Is a Different Problem

Inference workloads are a different problem. There are certainly inference workloads that require substantial east-west communication, especially as models become larger and serving becomes more distributed. But, generally speaking, inference infrastructure can have very different priorities from a massive training cluster.

For a network focused on inference workloads, request latency matters a lot. So do tail latency, tokens-per-second, the cost per token, and storage and memory movement. Also, traffic may be much more dynamic as demand rises and falls throughout the day.

Even NVIDIA’s own current GB300 reference architecture reflects this difference. Its guidance recommends the dedicated GPU east-west compute network for training, fine-tuning, machine learning, and HPC workloads, while listing that network as optional for pure inference.

So, training and inference may run on the same underlying GPU technology, but that doesn’t mean they should automatically receive identical network designs.

Agentic AI Changes It Again

Agentic AI adds another dimension.

An AI agent may not just send a prompt to a model and wait for a response. It may invoke several models, retrieve information from vector databases, interact with APIs, query enterprise systems, access persistent memory, call tools, and coordinate with other agents.

The infrastructure is therefore supporting a distributed application, not just a model.

AMD describes this in its current networking strategy: training requires tight synchronization across large numbers of GPUs, inference needs predictable latency under changing demand, while agentic workloads add persistent memory access and coordination among multiple services. That means the front-end network becomes very important in this architecture. 

An agentic AI factory may need exceptional GPU scale-out performance, but it may also need high-performance connectivity between GPUs, CPUs, storage systems, databases, inference services, APIs, security infrastructure, and external applications.

The architecture starts looking less like one giant GPU fabric and more like several highly specialized networks working together.

The AI Factory Is Really a Collection of Fabrics

This is probably a more useful way to think about an AI factory. Put very simply, it isn’t one network. Depending on the platform and workload, an AI factory can contain several distinct connectivity domains:

  • a GPU scale-up fabric inside a server or rack-scale system;

  • a GPU scale-out fabric connecting accelerators across systems and racks;

  • storage connectivity for datasets, checkpoints, model weights, and increasingly inference data;

  • a front-end or in-band network connecting compute to applications, customers, APIs, and other services;

  • an out-of-band management network;

  • and, for cloud operators, tenant and service networks that have to provide isolation between customers.

Not every AI factory needs those fabrics in exactly the same form. And the relative importance of each fabric changes according to what the AI factory is supposed to do.

A frontier training cluster may place enormous emphasis on GPU scale-out performance. 

An inference provider may care more about predictable latency, storage and memory movement, service delivery, and cost per token.

An enterprise agentic AI environment may create substantial traffic between inference infrastructure and conventional databases, applications, APIs, and storage.

A neocloud serving all three has to accommodate all of them while also providing secure multi-tenancy.

These each have have different design problems.

Reference Architectures Are a Starting Point, Not the End of the Conversation

None of this diminishes the importance of reference architectures. Especially for NVIDIA Cloud Partners, it’s very much the opposite.

AI infrastructure has become too complex and too expensive to casually improvise. Validated designs reduce integration risk and provide known-good combinations of compute, networking, storage, optics, software, power, and cooling.

NVIDIA explicitly describes its reference architectures as a way to package lessons learned from operating accelerated infrastructure at scale so customers and partners do not have to design everything from scratch.

That’s extremely valuable, but operators still need to understand what’s behind the reference architecture and how its assumptions relate to their environment.

  • What GPUs will be installed today?

  • What GPUs are likely to arrive two years from now?

  • Is the business selling dedicated training clusters, inference services, GPU instances, private AI environments, or some combination?

  • Will the infrastructure be single-tenant or multi-tenant?

  • How large are the scale-up domains?

  • What traffic crosses the scale-out network?

  • What does the storage architecture look like?

  • How quickly will the environment grow?

  • And what happens when the answers to those questions change?

Architecture Has to Follow the Business

For neoclouds and AI infrastructure providers, these are not just engineering decisions. They determine what the company can sell.

A provider selling dedicated frontier-training clusters has different infrastructure requirements from one selling serverless inference, and a sovereign AI provider has different constraints from a developer-focused GPU cloud. An enterprise AI factory running internal agents has different requirements from all three.

The infrastructure also has to survive multiple technology generations. Today’s environment might contain H100s, H200s, B200s, or GB300 systems. The next expansion might introduce Vera Rubin. Another part of the environment might use AMD Instinct GPUs. Acquisitions, customer requirements, supply constraints, and changing economics can make an originally homogeneous environment heterogeneous surprisingly quickly.

That makes operational flexibility an architectural requirement. The network needs to be designed for the AI factory the operator is likely to be running several years from now, not only for the hardware arriving on the loading dock today. 

Start With the Workload, Not the Diagram

The industry will continue to produce excellent reference architectures, faster switches, faster NICs, new accelerator interconnects, and increasingly integrated rack-scale systems. Operators should use them, but they should resist the idea that AI infrastructure has converged on one standard architecture.

GB300 and Rubin don’t have identical networking requirements, and NVIDIA and AMD ecosystems are not identical. Training, inference, and agentic AI create different traffic patterns and infrastructure priorities. Neocloud, sovereign AI, enterprise, and hyperscale operators have different business and operational requirements. The right architecture starts by understanding those differences.

So before asking which switch, NIC, topology, or reference architecture to deploy, we need to understand exactly what we’re building, what workloads will it run, how will it grow, and what does the network have to enable.

There may be reference architectures for AI factories, but there is no single reference AI factory.



Scott Robohn

Scott is co-founder and CEO of Solutional, where he leads initiatives in next-gen networking, automation, AI, and emerging technologies. With 35+ years of experience building, guiding, and scaling technical teams and solutions, Scott helps IT and NetOps organizations evolve into software-centric, resilient, and intelligent operations teams. His career has provided the depth and breadth of experience needed to lead technical sales organizations, including roles and partnerships with CTOs, Solutions Architects, Sales Engineers, Systems Engineers, Account Executives, Product and Engineering leaders, and other job functions. Scott is a frequent event speaker, host of the Total Network Operations podcast, and a co-founder of the Network Automation Forum (NAF).

Next
Next

Why GPU Clouds Are Not Just Smaller Public Clouds