Just a few decades ago, building a data center largely meant refining an established formula. Even as infrastructure evolved and facilities grew larger, the underlying architecture remained familiar.

Put simply, compute sat at the center, networking connected it together, and the physical layer evolved steadily alongside both. The rapid rise of AI has rewritten this familiar rulebook.

Beneath the surface of the most visible developments defining modern AI deployments – from GPUs, power and land availability, to increasingly advanced cooling systems – the architecture of the network itself is being reimagined.

“There is no one single AI architecture,” says Dr. Alan Keizer, senior technical advisor at AFL. “The complexity and variations are proliferating.”

Instead of one dominant blueprint that can be relied upon across multiple data center deployments, in its place is an increasingly diverse collection of specialized architectures.

Rather than simply assembling the best hardware, AI infrastructure demands applying the right architecture to the right workloads.

Against this backdrop, Keizer shares why designing smarter systems around the next-gen tech will increasingly define long-term success in the AI era.

A leakproof structure

The investment in the basic computing kit and switching that goes into modern data centers is huge. The prominence of these high-value components has naturally dominated conversation and shaped industry priorities. Fiber, by comparison, rarely commands the same attention.

“So where does fiber fit in?” Posits Keizer. “First of all, it's a relatively small portion of the capital investment on a greenfield build. Secondly, it’s like plumbing. It tends to be hidden.”

Rarely thought of when everything is working as expected, fiber’s value lies in quietly enabling everything else to perform.

“But like plumbing,” continues Keizer. “It’s critically important.”

While fiber may account for only a small proportion of the overall capital cost of an AI facility, it underpins communication between mission-critical infrastructure, enabling thousands of accelerators to operate as one coordinated system.

“It provides the connectivity that enables a network of high-performance and expensive kit to fulfil its full potential,” adds Keizer. “Any fault or problem immediately brings down the whole investment.”

Equally as important, fiber is one of the few parts of the infrastructure with the capacity to outlast the technology connected to it.

As technology cycles continue to develop at full speed, the fact that trunk and outside-plant fiber cabling can outlive multiple generations of GPUs represents a distinct strategic advantage, making the right network design decisions critical to ensuring scalability and operational longevity come plumbed in.

Building around the workload

The growing importance of selecting the right network architecture is also being driven by changing AI workloads. When AI adoption first took hold, attention centered around the enormous training clusters needed to develop frontier models. These training facilities quickly became synonymous with AI infrastructure – yet training represents only one part of a broader AI ecosystem.

Inference is quickly becoming a larger share of deployed AI capacity, while reasoning models and agentic applications represent increasingly viable ways of processing information. Each brings its own communication patterns, a unique balance of compute and memory, plus distinct architectural priorities.

As Keizer explains, much of the industry is still discovering what these new architectures should look like: “The textbook hasn’t been written yet. It’s being put together by multiple pioneers as we speak.”

Training remains the most demanding environment from a networking perspective. Developing a frontier model requires tens – and increasingly hundreds of thousands – of accelerators working together as one highly interconnected system.

“It’s a synchronous process,” explains Keizer. “Each GPU takes a small section of the work. When it’s done, it raises its hand, and when all the GPUs in the cluster have done their job, they all exchange data. This phase of calculation is only as fast as the slowest worker and the slowest network element.”

This is one of the defining differences between AI infrastructure and the enterprise networks that preceded it. In a traditional data center, the impact of one individual server slowing down was limited. Within a training cluster, every accelerator depends on every other accelerator progressing and working together. Network performance is therefore inseparable from compute performance – bringing with it significant physical consequences.

“We have several orders of magnitude more fiber coming in and out of GPU and switch racks than we ever had in enterprise or even traditional cloud data centers,” says Keizer. “There’s more count, higher density, and more criticality. Anything goes wrong, that section of compute is lost.”

Inference introduces a different set of priorities again. Rather than relying on huge synchronized clusters, many inference environments are organized into smaller pods sized to the model and

workload they serve. Connectivity within each pod is still performance sensitive, yet communication between pods resembles familiar cloud networking.

Reasoning models add further complexity. Instead of simply generating a response, a system may communicate with image generators, coding models, databases, or external services before landing on an answer.

Agentic AI takes this concept to the next level by maintaining context over longer periods while coordinating multiple tasks across different systems.

Fundamentally, this means tomorrow’s AI facilities are unlikely to be built around endless rows of identical GPU racks.

“We’re starting to see the design of inference-optimized facilities that are different in layout and structure than training-optimized facilities,” says Keizer. “The handwriting is absolutely on the wall.”

Infrastructure that evolves

Unlike accelerators, which are refreshed every few years, much of the physical infrastructure deployed today is expected to support multiple generations of hardware and a succession of different AI workloads.

For Keizer, the key to achieving this at the network level is recognizing that not every part of the fiber plant has the same lifespan:

“Let’s divide the fiber plant into four domains. First of all, what’s directly connected to the equipment. Next, what’s in the nearby fabric that connects the equipment. Then we have the trunk cabling moving fibers from zone to zone within the building. Finally, there are the external connections between buildings and across regions.”

These layers shouldn’t be designed with the same blanket approach. Equipment cabling changes alongside the equipment it serves. The fabric in the zone around it – the cabling that connects racks to each other – needs to flex as topologies shift, making minimizing connection counts so critical.

Trunk infrastructure within the building, along with campus connectivity and inter-data center links, should, therefore, be designed with a much longer horizon in mind.

“If designed correctly, long-haul and trunk cabling can live through multiple generations of equipment,” says Keizer. “It’s also hard to install and difficult to decommission, so it’s vital to understand what’s required for the long run.”

Building infrastructure capable of evolving without major disruption means looking beyond immediate capacity requirements by installing additional fibers within trunk routes, designing pathways that support future expansion, and creating layouts that can accommodate changing traffic patterns.

The physical layer as a strategic lever

As AI clusters continue to grow, many of the industry’s most pressing challenges are becoming more practical. Higher fiber counts demand higher connector densities; more accelerators require more structured cable routing; and larger campuses place greater pressure on pathways, ducts, and serviceability.

The fundamental challenge, therefore, is moving more data in a way that’s deployable and maintainable at scale.

“As these various kinds of equipment have become available, we’re seeing steadily more compute, storage, and switching going on in a given physical space,” says Keizer. “That means a lot of density.

“When I started out, a big number of fibers for a rack would be twenty-four. Now we’re talking about thousands per rack.”

Meeting these requirements has driven innovation in connector technology, cable design, and installation practices. Smaller form factors and efficient routing systems, for instance, are helping operators deliver more connectivity within the same physical footprint. Yet fitting more fiber into a rack is only part of the challenge.

“The real task lies in making it deployable at scale with reasonable economics and manufacturing velocity – while also being installable and maintainable,” explains Keizer.

As networks become more complex, the physical layer must deliver in terms of performance while also remaining serviceable throughout the facility’s lifetime. Accessibility, cable management, clear labelling, and accurate installation practices all become more critical as fiber counts move from thousands towards millions. Even something as routine as polarity takes on new significance.

Within hyperscale AI environments, where thousands of multi-fiber connections exist across highly complex fabrics, a simple inconsistency can become hugely difficult to identify.

“It’s increasingly critical to adopt a standard method,” adds Keizer. “Do it consistently, make it well documented, well labelled, and ensure everybody in the process understands it.”

At AI scale, attention to detail increasingly determines how efficiently and effectively facilities can be deployed, maintained, and scaled over time.

The architecture reveals itself

For operators taking on new AI deployments, the most important decisions may not come when selecting GPUs or specifying hardware. In today’s landscape, they’re made much earlier – when defining the purpose of the facility itself.

“The first and most important question is what kind of AI do you anticipate doing?” Says Keizer. “All the decisions that follow will be determined by this fact, and you’ll adopt a completely different approach doing Edge inference than if you’re going to try to create the next great model in a training facility.”

The workload determines the communication patterns, which shape the network topology, which in turn defines the physical layout and the fiber plant, while contributing to the long-term flexibility of the entire facility.

The blueprint for AI infrastructure must be drawn by the workload itself, supported by a physical layer designed to evolve with it, and underpinned by an architecture capable of flexing to whatever lies ahead.

Read the full white paper here.