There is a narrative that speaks of AI factories with reverence once reserved for steel mills, as gigawatt campuses producing intelligence as a commodity. It’s a compelling vision, and one the industry has embraced with fervor. The hyperscalers have billions committed, and the race to secure nuclear power agreements all reinforce a singular narrative: bigger is better, and the future belongs to those who build at unprecedented scale.
But this industrial metaphor, while useful for training infrastructure, obscures a more complex reality emerging in inference. The workloads that generate revenue, the queries answered, the agents deployed, and the decisions automated simply don’t share training’s appetite for centralization. They actively resist it.
The physics of impatience
Agentic AI changes the latency calculus entirely. When a large language model responds to a single query, 200 milliseconds of network delay is imperceptible. But when AI agents communicate with other AI agents – orchestrating workflows, querying databases, executing multi-step reasoning chains – that tolerance evaporates.
Industry leaders in AI networking have acknowledged this shift. Sub-millisecond latencies become essential when agentic workflows compound round-trips across inference calls. A reasoning chain that pings between coasts a dozen times accumulates delay that degrades performance measurably. The speed of light, it turns out, remains undefeated.
This creates natural geographic constraints that no amount of capital expenditure can overcome. Agentic inference wants to live close to the systems it orchestrates and the data it reasons over. That’s not a 3GW campus in a rural location – it’s a distributed presence across metropolitan cores, enterprise data centers, and network edge locations where latency budgets can actually be met.
The great bifurcation
The industry is slowly recognizing what the workloads have already decided: training and inference are diverging into fundamentally distinct infrastructure classes with incompatible requirements.
I wrestled in high school and they have weight classes for a very specific reason – people are different sizes and size matters. You wouldn’t put a big guy and a small one on the same mat and expect a fair match. Training and inference infrastructure are splitting into their own weight classes, and forcing them to share the same facilities makes about as much sense as expecting a 200-pound heavyweight to move with a 108-pounder’s speed.
Training clusters optimize for throughput and scale. They tolerate – even prefer – remote massive locations where power and water are cheap and abundant. Latency to end users is irrelevant when you’re processing petabytes of training data over weeks or months. These are the true AI factories: industrial facilities that can operate far from population centers, drawing hundreds of megawatts from dedicated substations or behind-the-meter generation.
Inference infrastructure inverts nearly every assumption. It optimizes for responsiveness, not throughput. It follows users and data rather than power availability. It operates in bursts tied to demand patterns rather than sustained maximum utilization. And critically, it deploys smaller, often distilled models that sacrifice some capability for dramatic improvements in efficiency and speed.
Industry analysts now forecast that global investment in AI inference infrastructure has surpassed training infrastructure – a crossing point that signals where value is migrating. A new class of inference-focused providers is responding accordingly, developing platforms specifically targeting inference workloads with token-based pricing. Some process trillions of tokens daily, while others pursue custom silicon delivering significantly higher token throughput than GPU-based alternatives for the smaller, distilled models that dominate production deployments.
Operators planning for a future where training and inference share common infrastructure may find themselves optimized for neither.
Communities have veto power
There’s another force fragmenting the AI factory vision that is just appearing in infrastructure funding documents: communities have learned to say no – loudly and effectively.
The data center industry’s social license is eroding faster than most executives acknowledge. According to industry tracking sources, billions of dollars in projects were blocked or delayed in early-to-mid 2025 alone. Opposition has surged dramatically year-over-year, with a growing number of activist groups now organized across the United States.
The geographic spread tells the story. In recent months, several major hyperscaler and technology company data center proposals have been withdrawn, rejected, or indefinitely delayed following organized community resistance. Concerns have ranged from water consumption and environmental impact to grid strain and insufficient local employment benefits. In multiple cases, projects were halted even after significant investment in site acquisition and planning.
This opposition is not confined to any particular region or political orientation. Concerns about data center development span the political spectrum: some officials focus on the adequacy of tax incentives and grid reliability, while others emphasize environmental impacts and resource consumption. The result is a broad-based movement that infrastructure planners can no longer treat as a fringe concern.
Gigawatt-scale AI training campuses amplify every tension point. They demand grid interconnections that take years to permit. They concentrate economic disruption in communities that bear environmental costs without proportional employment benefits. They become visible symbols of an industry that – fairly or not – is increasingly associated with excess.
Distributed inference infrastructure offers a different political economy. Smaller facilities disperse impact across more jurisdictions. They integrate into existing commercial and industrial zones rather than requiring greenfield development. They can participate in local grid services, providing demand response and frequency regulation that positions them as community assets rather than extractive burdens.
This isn’t merely about public relations. Permitting timelines increasingly determine who can deploy capacity when it matters. Operators who can site numerous modest-scale inference nodes across dozens of locations may bring more aggregate capacity online faster than those fighting multi-year battles for single massive developments.
Who leads the transformation?
If inference fragments into a distributed topology, a critical question emerges: who is best positioned to lead?
The hyperscalers have obvious advantages – global footprints, existing edge presence through CDN infrastructure, and deep pockets. But their architectures tend to optimize for their own service ecosystems, which may not fully address the diverse requirements of enterprise AI deployments. The gap between what hyperscalers currently offer and what distributed inference demands creates opportunity for new categories of providers.
A class of inference-native platforms is moving fastest. These companies have built entire businesses around optimizing model serving, not GPU rentals. Some focus specifically on the software layer that routes, manages, and accelerates inference workloads. Others are explicitly targeting metro edge deployment, recognizing that inference belongs close to users, not in remote mega-campuses.
Traditional colocation providers may hold underappreciated advantages. The major colo operators already run distributed infrastructure in the metropolitan cores where inference latency matters most. They understand permitting, community relations, and grid interconnection at a granular, local level that builders of greenfield campuses cannot easily replicate. The colo operators who can offer inference-ready capacity – liquid cooling, high-density power, low-latency interconnection – across dozens of metropolitan markets simultaneously may capture value that eludes both the mega-campus builders and the GPU-focused neoclouds.
Telecom operators, often dismissed as legacy infrastructure, are quietly positioning for relevance. Some now offer GPU-as-a-Service through strategic partnerships, while others are building national AI infrastructure with explicit attention to local language models and regional requirements. These players bring something the neoclouds often lack: existing physical presence in every community, established regulatory relationships, and infrastructure that already reaches the edge.
Orchestration becomes the high ground
Regardless of who owns the physical infrastructure, the strategic question shifts from “who has the most GPUs” to “who can orchestrate workloads across heterogeneous infrastructure most effectively.”
Wrestling taught me early that strength alone wouldn’t win matches – leverage and position matter. If I could maintain inside position and a low center of gravity I could beat a stronger opponent. The same principle applies to distributed inference: the winners won’t be those with the most GPUs, but likely those who position workloads where physics and economics give them leverage.
This is the layer where value should concentrate. The orchestration challenge spans multiple dimensions simultaneously: routing queries to nodes that meet latency requirements, balancing load across facilities with different cost structures, maintaining model consistency across distributed deployments, and gracefully degrading service when individual nodes fail or become congested.
Today’s neocloud providers compete primarily on GPU access – who can offer next-generation accelerator capacity at what price point. Tomorrow’s winners will likely compete on intelligence: the ability to abstract away infrastructure complexity and deliver inference as a seamless service regardless of where computation actually occurs. Recent major investments in both specialized inference silicon and orchestration software reveal an industry betting that both the hardware and the software layers of distributed inference represent strategic control points.
Some will attempt to build walled gardens – proprietary orchestration that locks customers into specific infrastructure. History suggests this approach struggles against open alternatives that give enterprises flexibility to match workloads with optimal execution environments. The orchestration layer that wins will likely be the one that embraces heterogeneity rather than fighting it.
The factory and the network
None of this suggests mega-scale training facilities lack a future. The frontier models that define AI capabilities will continue emerging from massive centralized clusters. The AI factory metaphor remains apt for that domain. The point is, it’s important to distinguish the demand and requirements for both classes of data centers.
But the industry’s current investment thesis – that training infrastructure naturally extends to inference at similar scale – may deserve some scrutiny. The workloads are different. The economics are different. The physics are different. And the communities that must accept this infrastructure have demonstrated they expect a voice in these decisions.
The AI factory produces the models. The inference archipelago deploys them. Both are essential, but they will increasingly develop along separate evolutionary paths, with distinct operators, financing structures, and geographic footprints.
Those building for the next decade should plan accordingly. The biggest clusters will train the smartest models. But the networks that deliver intelligence to where it’s needed – orchestrated across a fragmented topology of specialized nodes – may ultimately prove more valuable than the factories themselves.
Comments