AI is changing data-centre design because accelerated computing can increase rack density, network traffic, cooling requirements, and the importance of high utilization. The scale of the build-out is real, but the most useful questions are operational: how much capacity is needed, where power and grid connections are available, how efficiently equipment will be used, and what reliability, water, carbon, and community constraints apply.

Start with measured use and explicit scenarios

The International Energy Agency estimates that data centres consumed about 415 terawatt-hours of electricity worldwide in 2024, roughly 1.5% of global electricity consumption. In its base case, the IEA projects consumption of around 945 TWh in 2030, just under 3% of the global total. Accelerated servers, driven mainly by AI adoption, account for almost half of the projected net increase.

The geographic concentration matters more than the global percentage alone. The IEA notes that data-centre capacity is clustered and that grid-connection constraints could delay planned projects. A facility can therefore have major local effects on generation, transmission, distribution, water systems, and other customers even when the sector remains a modest share of global electricity.

For the United States, Lawrence Berkeley National Laboratory estimates 176 TWh of data-centre electricity consumption in 2023, equal to 4.4% of U.S. electricity. Its 2028 scenarios range from approximately 325 to 580 TWh, or 6.7% to 12.0% of projected U.S. electricity consumption. The range is wide because accelerator shipments, utilization, operating power, and cooling designs are uncertain. The report also warns that limited public data constrains the analysis.

A spending forecast is not a market fact

Investment forecasts can be useful, but only when their boundaries are explicit. Some count servers and networking; others include buildings, land, grid upgrades, generation, or long-term capacity commitments. Some report one year of capital expenditure; others sum several years. None should be presented as guaranteed revenue or economic value.

When evaluating a forecast, record:

  • the publisher and date;
  • the geography and forecast horizon;
  • whether the figure is annual or cumulative;
  • which assets and workloads are included;
  • the demand, utilization, efficiency, financing, and power-price assumptions; and
  • the downside and sensitivity cases.

What changes in an AI-oriented facility

Compute and memory

Accelerators are selected as a system, not by brand label alone. Training and inference requirements differ, and useful comparisons include numerical precision, accelerator memory, memory bandwidth, interconnect topology, model size, batch and latency targets, software compatibility, utilization, power draw, and failure recovery.

Networks and storage

Distributed training can require high-bandwidth, low-latency east-west networking. Data ingestion, checkpoints, model artifacts, and evaluation traces can stress storage in different ways. Capacity planning should use measured workload profiles and end-to-end tests rather than assuming that NVMe, a particular fabric, or a container orchestrator automatically removes bottlenecks.

Power and cooling

Higher rack densities can change power distribution and cooling design. Liquid cooling may be appropriate for dense equipment, but the environmental outcome is site-specific. Operators should measure IT energy, facility overhead, water consumed on site, water associated with electricity generation, grid carbon intensity, backup generation, refrigerants, and embodied impacts.

Scheduling and utilization

Idle accelerators still consume energy and tie up scarce capital. Good scheduling, rightsizing, checkpointing, queue policies, workload placement, and observability can improve useful work per unit of infrastructure. “AI-defined infrastructure” is not a standard architecture; automation should remain bounded by tested policies, safety constraints, change control, and rollback.

What the Google cooling example establishes

In 2016, Google DeepMind reported applying neural networks to sensor data and control recommendations in a live Google data centre. It reported a 40% reduction in energy used for cooling, corresponding to a 15% reduction in overall PUE overhead after other losses and non-cooling inefficiencies. That is a notable organization-specific result. It is not evidence that every facility will cut total energy or cost by 40%, and it should not be combined with invented dollar savings.

Sustainability requires system boundaries

The IEA projects renewables to meet nearly half of the growth in data-centre electricity demand through 2030 in its base case, while natural gas and coal together meet more than 40% of the additional demand over that period. Renewable-energy contracts are important, but annual matching does not necessarily mean a facility is powered by carbon-free electricity in every location and hour. Claims should distinguish physical grid supply, contractual matching, and 24/7 carbon-free-energy goals.

Microsoft’s Project Natick demonstrated a renewable-powered subsea data-centre research module and reported encouraging Phase 2 reliability results. The vessel was retrieved in 2020. It should be described as a concluded experiment, not an active default solution for data-centre cooling.

Security and resilience priorities

  • maintain an inventory of hardware, firmware, software, models, data, identities, and dependencies;
  • segment management, storage, training, inference, and tenant networks;
  • use least privilege, strong workload identity, secrets management, and auditable administrative access;
  • verify accelerator, firmware, container, model, and data supply chains;
  • monitor infrastructure and application telemetry with tested detection and response procedures;
  • protect training data, model weights, checkpoints, and inference logs according to their sensitivity; and
  • test capacity failures, cooling and power events, regional dependencies, restoration, and rollback.

A decision checklist

  1. Forecast demand as a range and separate training, batch inference, and latency-sensitive inference.
  2. Model utilization and useful work, not only nameplate accelerator capacity.
  3. Confirm utility interconnection, generation, transmission, backup, and curtailment assumptions.
  4. Evaluate PUE, WUE, carbon intensity, water stress, and lifecycle impacts at the proposed location.
  5. Test performance and failure behavior on representative workloads before committing to scale.
  6. Plan for modular growth, heterogeneous hardware, model-efficiency gains, demand uncertainty, and asset stranding.
  7. Publish the assumptions behind financial and sustainability claims and update them as measured data arrives.

Bottom line

AI is driving a rapid and uncertain infrastructure expansion. The credible opportunity is not a single headline number; it is the work of delivering useful computation within power, grid, reliability, security, water, carbon, and financial constraints. Plans should be built from measured workloads and transparent scenarios, then revised as utilization and resource data become available.

Primary sources