AI Infrastructure Watch: The AI Race Becomes a Systems Race

Written by Erwin Castro — Founder & Editor, The CODEW
AI Infrastructure Watch | September 18, 2026

The AI Infrastructure Race Is Becoming a Systems Race

AI Infrastructure Watch | September 18, 2026 cover

Executive Brief

The Scale of the Buildout

The first phase of AI infrastructure was deceptively simple: buy GPUs → build data centers → train larger models. The second phase is proving far more complicated.

The Bank for International Settlements estimates that the five largest global technology companies will invest more than $1 trillion in AI during 2025–2026. Alphabet has raised its 2026 capital spending outlook to $195–$205 billion, while Microsoft plans to more than triple its data-center footprint to about 38 gigawatts by 2032. TrendForce projects global data-center electricity demand capacity will reach 161 GW in 2026, up roughly 31% year over year, with AI servers accounting for 33.4% of that total.

But the central shift is not about scale alone.   As AI moves toward agentic and reasoning workloads, the stack increasingly looks like: Compute → Memory → Networking → Data Centers → Power → Cooling → Infrastructure Software → Inference.

Compute Is Still the Foundation

AI compute continues to expand aggressively. Nvidia's Grace Blackwell shipments grew 27% month-over-month, with Nvidia CEO Jensen Huang confirming that OpenAI's GPT-6 Astra was trained on over 100,000 Grace Blackwell NVLink72 chips. Crusoe, the vertically integrated AI infrastructure provider, raised $3.9 billion in Series F funding at a $30.9 billion valuation, with over $140 billion in total contracted value across its platform and 6 GW of gross contracted capacity.

But the compute question is no longer simply "how many GPUs?" Custom silicon is gaining ground rapidly. Meta is expected to begin production of its latest custom AI chip, while Qualcomm granted Amazon warrants tied to up to $60 billion in chip purchases over the next decade for AI inference semiconductors. The inference era is changing what compute infrastructure must deliver.

The Inference Era Changes Infrastructure

Training dominated the early AI infrastructure narrative. Now inference is becoming the primary driver of distributed infrastructure investment.

Reasoning and agentic systems generate fundamentally different workload patterns. A single user request can trigger: reasoning → tool call → result → additional reasoning → action → verification. This creates repeated, bursty inference workloads that demand low latency, high throughput, and efficient memory bandwidth—not just raw compute performance.

DeepSeek is reportedly planning to deploy at least 160,000 Huawei Ascend 950DT chips for inference in Inner Mongolia, with the project targeting gigawatt-scale power capacity. Axelera AI is shipping its Europa AI Processing Unit, delivering 629 TOPS at 45W for on-premises inference in validated Dell and Supermicro servers, targeting enterprises that need to keep data on-premises.

The infrastructure implication is clear: inference is not simply training at smaller scale. It requires different optimization targets—cost per task, latency, model routing, and GPU utilization—that are reshaping how AI infrastructure is designed and deployed.

Memory Is Becoming a Critical Constraint

The HBM shortage has become the defining bottleneck of AI infrastructure. Samsung and SK Hynix inventories have fallen to less than 10 days' worth, according to KB Securities, as AI infrastructure spending drives demand for HBM, server DRAM, and enterprise SSDs. The transition to HBM4—which requires roughly three times the wafer capacity of conventional DRAM—is adding severe pressure.

KB Securities estimates DRAM and NAND bit-demand growth will outpace supply by more than 10% in 2027, potentially creating the tightest memory supply on record. Memory is estimated to account for 57% of total AI infrastructure investment next year, up from 14% last year, according to KB Securities, with TrendForce projecting the share could reach as high as 68%.

The critical question: Can AI compute scale if memory cannot scale alongside it? For now, the answer appears to be no—at least not at the pace hyperscalers require.

Networking Becomes AI Infrastructure

Thousands of accelerators are only useful if they can communicate efficiently. Networking has become a first-class infrastructure concern.

Ethernet has overtaken InfiniBand in AI back-end networking. The Ultra Ethernet Consortium released Specification 1.0 in mid-2025, standardizing how Ethernet matches InfiniBand's lossless characteristics for AI workloads. Now a second wave is arriving: the ESUN (Ethernet for Scale-Up Networking) standard, launched within the Open Compute Project, targets GPU-to-GPU communication within a rack—the domain previously dominated by Nvidia's proprietary NVLink. Founding members include AMD, Arista Networks, Arm Holdings, Broadcom, Cisco, HPE, Marvell, Meta, Microsoft, OpenAI, Oracle, and Nvidia itself.

The strategic significance is clear: AI infrastructure requires a networking fabric that can scale alongside compute, and the industry is actively building alternatives to single-vendor dependency.

Power Becomes the Hard Limit

The AI infrastructure race is increasingly becoming a race for available electricity. TrendForce estimates that the gap between what power grids can deliver and what data centers demand will widen dramatically after 2028, reaching a 268 GW shortfall by 2030.

Nvidia has committed $2 billion to Brookfield's AI infrastructure fund, which is focusing on backing factories, dedicated behind-the-meter power solutions, and compute infrastructure. Crusoe's energy-first strategy—originating and managing power before site selection—reflects a broader recognition that power availability is now the binding constraint on AI infrastructure deployment.

Cooling and Data-Center Architecture

The physical challenges of AI infrastructure are intensifying. TrendForce estimates liquid cooling penetration in AI chips will rise from 33% in 2025 to 53% in 2026, approaching 60% by 2027. Nvidia's Vera Rubin platform has abandoned traditional air cooling entirely, adopting a fanless, fully liquid-cooled architecture.

New solutions are emerging at scale: Ningchang Information launched the world's first MW-class single-phase cold plate fully liquid-cooled supernode, supporting up to 5,000W per chip and 28 kW per 1U node. The question is no longer whether liquid cooling will be adopted—it is whether data-center architecture can evolve fast enough to support the next generation of AI hardware.

Infrastructure Software Becomes the Control Layer

As AI clusters grow larger and more heterogeneous, infrastructure software is becoming the control plane that coordinates compute, storage, networking, scheduling, and workloads.

Sharon AI signed a five-year agreement with Rafay Systems to deploy a centralized orchestration and operations layer across its AI Factory environments, designed to support up to 150,000 GPUs. The Rafay platform provides automation, observability, governance, and multi-tenancy capabilities across different locations and workloads. Orchestra launched its Agentic Control Plane, a platform that unifies data pipelines and AI agents in a single environment, claiming up to an 80% reduction in pipeline execution costs and a 95% reduction in development time.

The infrastructure stack is increasingly becoming hardware + software + automation—and the software layer is where operational scale is won or lost.

The AI Infrastructure Stack Is Reorganizing

The stack is converging:

Compute → Memory → Networking → Data Centers → Power & Cooling → Infrastructure Software → Inference

Bottlenecks can migrate from one layer to another. A shortage of GPUs may eventually be replaced by shortages in power, HBM, networking, data-center capacity, grid connections, or cooling. The next AI bottleneck may not be intelligence. It may be infrastructure.

The CODEW Take

AI's next bottleneck may not be intelligence. It may be infrastructure. The companies that control the scarce layers—memory, power, networking, and orchestration—will define how fast the AI economy can scale. The winners will be those who treat AI infrastructure as an integrated system, not a collection of components.

What to Watch Next

  1. AI inference growth
  2. GPU and accelerator capacity
  3. HBM availability
  4. AI networking
  5. Data-center construction
  6. Power procurement
  7. Liquid cooling
  8. Neocloud expansion
  9. Custom AI infrastructure
  10. Infrastructure software automation

The CODEW Stat

The five largest global technology companies will invest more than $1 trillion in AI during 2025–2026, while global data-center electricity demand capacity reaches 161 GW in 2026—up 31% year over year—with AI servers accounting for 33.4% of the total.


Sources

  1. Zacks Analyst Blog — AI Capex Creates New Investment Cycle (September 14, 2026)
  2. Bloomberg — Nvidia Committed $2 Billion to Brookfield AI Infrastructure Fund (September 17, 2026)
  3. Crusoe — Crusoe Raises $3.9 Billion Series F (September 17, 2026)
  4. Digitimes — Weekly news roundup: HBM4 strain, Intel price hikes (September 14, 2026)
  5. Nasdaq — Sharon AI Selects Rafay Systems (September 4, 2026)
  6. Le Monde Informatique — Orchestra lance son Agentic Control Plane (September 3, 2026)



Editorial Note

AI Infrastructure Watch is The CODEW's weekly intelligence product tracking the physical, computational, and technical infrastructure required to make AI work at scale. Coverage spans compute, memory, networking, data centers, power, cooling, infrastructure software, and inference—connecting the layers that determine how fast the AI economy can grow.


AI Infrastructure Watch: The AI Race Becomes a Systems Race AI Infrastructure Watch: The AI Race Becomes a Systems Race Reviewed by Erwin Castro on Friday, September 18, 2026 Rating: 5
CRM + marketing automation + payments in one integrated platform. Helps small businesses streamline sales and automate the follow-up work that falls through the cracks. Get Keap