AI infrastructure: Data speed is now the bottleneck for chip performance

A single high-end GPU can run between $20,000 and $30,000

The artificial intelligence boom, so far, has been defined by a single instinct: build more. More chips, more data centres, more power. Capital expenditure from the world’s largest tech firms is running into the hundreds of billions as they race to scale AI systems.

Analysts expect global data centre investment to continue rising sharply, with major cloud providers alone guiding tens of billions in annual capital expenditure tied directly to AI infrastructure.

New projects, or Elon Musk’s proposed ‘Terafab’ chip complex in Texas, prove that demand for compute will keep outstripping supply.

The Data Bottleneck: Why Raw Compute Isn’t Enough

Despite massive investment in processing power, a critical inefficiency is emerging. Jürgen Hatheier, Ciena’s vice president of business development, points out that even highly optimized data centres experience GPUs spending 50% of their time waiting for data. A single high-end GPU can cost between $20,000 and $30,000 and consume around a kilowatt of power. This idle time translates into significant financial and energy waste.

The Connectivity Challenge in the AI Era

The pace of innovation in compute power is outpacing advancements in connectivity. As Hatheier explains, compute innovation has been happening three times faster than connectivity innovation. This mismatch is particularly problematic as AI workloads scale, requiring vast datasets to move between machines and data centres. Datasets can reach tens of petabytes, meaning even high-capacity networks can take days or weeks to transfer information without optimization.

Beyond Training: The Demands of AI Inference

The challenges extend beyond training large language models to the day-to-day operation of AI tools – inference. If networks can’t keep up with inference demands, machines pause, further exacerbating the problem. The rise of agentic AI systems, where multiple AI agents coordinate tasks, intensifies this need for rapid data exchange. Delays in data handling can account for the majority of system latency, leaving expensive compute resources underutilized.

The Shift in Infrastructure Priorities

Simply owning the best chips is no longer sufficient. The speed at which those chips can be fed data and how efficiently they can work together, is now paramount. This is driving investment in fibre routes and pre-deployed infrastructure as competitive advantages. Operators are investing ahead of demand, anticipating future AI workloads.

The Broader Supply Chain Under Pressure

The strain isn’t limited to networking. Electricity infrastructure in the US is struggling to keep pace with data centre demand, facing equipment shortages and rising costs. Semiconductor production also faces bottlenecks, with advanced manufacturing capacity falling short of projected AI demand. Enterprises are increasingly experimenting with their own on-premise GPU clusters, adding further pressure on networks.

The CPU’s Resurgent Role

Alongside connectivity, the role of CPUs in coordinating AI systems is becoming increasingly important. New workloads are driving demand for server processors to handle data movement and task execution, with estimates suggesting millions of CPUs will be required for next-generation AI deployments. Supply is struggling to meet this demand, with chipmakers warning of shortages and extended lead times.

Future Trends: What to Expect

The current situation points to a system under significant strain. Future developments will likely focus on:

  • Advanced Networking Technologies: Expect further innovation in fibre optics, RDMA (Remote Direct Memory Access), and other technologies designed to minimize latency and maximize bandwidth.
  • Co-location and Edge Computing: Bringing compute closer to the data source will reduce the need for long-distance data transfers.
  • Optimized Data Management: More sophisticated data compression, caching, and pre-fetching techniques will be crucial.
  • Integrated Hardware and Software Solutions: Closer integration between hardware and software will allow for more efficient data flow and resource allocation.

Leave a Comment