Data Center Engineering
See also (Tier 3 family indexes): standards-bodies, energy-storage-systems, battery-chemistries, refrigerants, heat-transfer-correlations.
1. At a glance
A data center is a purpose-built facility that houses computing infrastructure — racked servers, switches, storage, and the supporting power, cooling, fire-protection, and connectivity systems that keep them online. The discipline of data center engineering sits across transformers-power-systems, hvac-fundamentals, heat-transfer, fluid-mechanics, digital-control (BMS / DCIM), structural-analysis (slab loading, seismic anchoring of racks), and reliability-engineering (concurrent-maintainability, MTBF modeling). Modern hyperscale and AI-training facilities couple all of these into a single integrated design problem at the gigawatt scale.
The hardware spans seven orders of magnitude:
- A 4 kW server closet in a small office, with one 5 kVA UPS and a wall-mount split system.
- A 50 kW retail “edge” pod at a cell tower or factory floor — single-row containment, in-row cooler, lithium-ion UPS.
- A 1 MW enterprise on-prem data hall, two N+1 chillers, dual utility feeds, generator backup.
- A 10–30 MW colocation suite at Equinix, Digital Realty, NTT, or CyrusOne — multi-tenant, customer-cage architecture.
- A 100–500 MW hyperscale building (one of many on a single campus) — AWS, Azure, Google, Meta, Apple, Oracle.
- A 1+ GW AI training campus — multi-building, dedicated substation, sometimes behind-the-meter generation. Microsoft + OpenAI’s “Stargate” series, xAI Colossus (Memphis), Meta Hyperion (Louisiana), Google TPU pods.
- A theoretical 5 GW campus — the upper limit announced for 2027–2030 by several hyperscalers, beyond which transmission and water become binding constraints.
The unifying engineering problem: deliver clean, redundant power to thousands of racks; remove the heat at near 100 % efficiency; connect everything at low latency; survive component failures without dropping a workload; do all of this at a Power Usage Effectiveness (PUE) close to 1.10. AI workloads add a new constraint: per-rack power has risen from 5–10 kW (2015) to 60–120 kW (2024–2025), forcing the air-to-liquid cooling transition.
2. Why it matters
Global data center electricity demand was ~460 TWh in 2022 (IEA), ~1.5 % of global consumption. The IEA’s 2024 “Electricity 2024” outlook projects 620–1050 TWh by 2026, then \$1 trillion+ of announced AI-infrastructure capex in 2024–2026 alone is expected to push that past 1500 TWh by 2030 — roughly the entire electricity consumption of Japan. Single hyperscale campuses are now larger than entire metro-area power loads were a decade ago: xAI Colossus crossed 200 MW in 2024 with a public roadmap to 1.2 GW.
The economic stakes:
- Hyperscaler capex in 2024 was \320–360B. Roughly half flows to land, power, mechanical, and electrical (the “shell + base build”), the rest to IT.
- Cost of downtime at a Tier IV hyperscale facility is conservatively \25k per minute of outage at the workload level; at the customer revenue level (a brokerage exchange, a payment processor) it can hit \$10M per hour.
- PUE delta of 0.05 at a 100 MW hyperscale = ~44 GWh/year of avoided electricity = ~\0.07/kWh.
- Cooling-water consumption at evaporative-cooled hyperscale facilities is now politically salient — Google’s 2023 environmental report disclosed 5.6 billion gallons of fresh water consumption, driving a public retreat toward closed-loop and dry coolers in water-stressed regions.
The other half: AI-training cluster availability is now a competitive moat. A 100k-GPU cluster that drops 2 % of training steps to unrelated infrastructure events loses ~\$5M in compute-equivalent productivity per week. Reliability has moved from an SRE concern to a capex-justified design driver.
3. First principles — the four parallel systems
A data center is the interaction of four parallel utility systems plus the IT load they serve:
- Electrical — utility → switchgear → UPS → distribution → rack PDU → server PSU. Must survive utility loss, transformer failure, switchgear maintenance, breaker trip, and PDU failure without taking down the IT load.
- Mechanical (cooling) — chiller / dry cooler / cooling tower → CHW or condenser-water loop → CRAH / in-row / rear-door / direct-to-chip → silicon junction. Must reject 100 % of IT power as heat plus the cooling-plant inefficiency.
- Network — internet peering → border routers → spine → leaf → top-of-rack → server NIC. Carries the workload data plus storage traffic plus the management / out-of-band channel.
- Mechanical (auxiliary) — fire detection (VESDA aspirating smoke), fire suppression (pre-action sprinkler + clean-agent for some rooms), make-up water, humidification, leak detection, security (mantraps, biometrics).
Each of these is sized to the IT load (commonly expressed in kW per rack × number of racks + an overhead factor) and then duplicated to the redundancy level dictated by the facility tier. The cleanest abstraction is to treat each as a single-line diagram and trace any failure point to its impact on IT continuity.
3.1 The PUE identity
The Power Usage Effectiveness, defined by The Green Grid in 2007 and codified in ISO/IEC 30134-2:2016, is:
PUE = Total Facility Energy / IT Equipment Energy
PUE = 1.0 is theoretical perfection (zero overhead). Practical floors:
| Facility class | Typical PUE | Notes |
|---|---|---|
| Hyperscale, climate-favored, modern | 1.08–1.15 | Google’s fleet 1.10, Meta 1.09, AWS 1.12 (FY2024 disclosures) |
| Hyperscale, hot/humid climate | 1.20–1.30 | Phoenix, Singapore, Dubai with chillers running year-round |
| Modern colocation | 1.30–1.50 | Customer-driven racks make airflow management imperfect |
| Enterprise on-prem | 1.50–1.80 | Legacy CRAC + leaky containment |
| Legacy room without containment | 2.0+ | Often dominated by infiltrating cold air bypassing the load |
PUE is a snapshot — it varies with outdoor air temperature, IT load level, and operating mode. Modern hyperscale reports a trailing twelve-month PUE plus a design PUE at design point.
Related metrics:
- WUE Water Usage Effectiveness (L/kWh of IT) — ISO/IEC 30134-9. Evaporative-cooled sites often 1.5–2.0 L/kWh; closed-loop dry-cooled designs <0.1.
- CUE Carbon Usage Effectiveness (kg CO₂e/kWh of IT) — ISO/IEC 30134-8. Driven by grid emissions factor + on-site generation.
- REF Renewable Energy Factor — fraction of consumption matched by renewables (24/7 carbon-free matching being the new bar).
- ERE Energy Reuse Effectiveness — credit for waste-heat reuse, used in district-heating-coupled European DCs.
4. Facility taxonomy
4.1 Hyperscale cloud
Operated by the cloud providers themselves for first-party + tenant workloads. Standardized building blocks (“colos,” “data halls,” “fabric pods”) replicated across campuses. Power densities 12–40 kW/rack standard, AI halls 60–120 kW/rack with direct-to-chip liquid and 130–250 kW/rack with rear-door + immersion combinations.
| Operator | Notable campus(es) | Notes |
|---|---|---|
| AWS | Northern Virginia (Loudoun + Prince William counties, ~70 buildings), Oregon (Umatilla), Ohio (Hilliard), Dublin, Frankfurt | The Loudoun cluster carries ~70 % of US-East-1 traffic; ~13 GW disclosed pipeline 2025 |
| Microsoft Azure | Wyoming (Cheyenne), Quincy WA, San Antonio TX, Goodyear AZ, Dublin, Amsterdam, Mt. Pleasant WI | Microsoft + OpenAI Stargate sites announced 2024 (Wisconsin, Texas, Abilene TX as first); target 5+ GW each |
| Google Cloud | Council Bluffs IA, The Dalles OR, Mayes County OK, Henderson NV, Eemshaven NL, Hamina FI | Free-cooling Hamina pumps Baltic seawater; Eemshaven uses Dutch grid wind |
| Meta | Prineville OR, Forest City NC, Altoona IA, Eagle Mountain UT, Clonee IE, Luleå SE, Richland Parish LA (Hyperion, 2 GW announced 2024) | 100 % wet-bulb evaporative + outside-air economizer; pioneered Open Compute Project (OCP) |
| Apple | Maiden NC, Reno NV, Mesa AZ, Prineville OR, Viborg DK | Lower disclosed footprint; significant capacity at Equinix and Google Cloud |
| Oracle (OCI) | Ashburn VA, San Jose CA, Frankfurt, Phoenix, Hyderabad | OCI Region buildout 50+ regions; growing AI cloud footprint |
| xAI | Memphis “Colossus” | 100k H100 GPUs Q4 2024 → 200k+ Q1 2025; 250 MW disclosed, behind-the-meter natural-gas turbines while utility transmission catches up |
| Tencent / Alibaba / Baidu / ByteDance | Guizhou, Inner Mongolia, Hebei, Singapore | Chinese hyperscalers driven by national “East Data West Compute” plan locating compute in low-cost-power western provinces |
4.2 Colocation
Multi-tenant facilities — landlord owns the building, power, and cooling, tenants own the IT. Tenants buy by the cabinet, cage, or suite, with metered or contracted power.
| Operator | Footprint |
|---|---|
| Equinix | 270+ “IBX” sites across 70+ metros; the dominant interconnection player (90 % of internet routes traverse one of its sites) |
| Digital Realty / Interxion | 300+ sites; merged Interxion 2020 for European depth |
| NTT Global Data Centers | 160+ sites; strong APAC + India presence |
| CyrusOne (KKR + GIP, taken private 2022) | ~50 sites; US + Europe wholesale focus |
| CoreSite (American Tower) | ~25 sites; major US carrier-hotel positions |
| QTS (Blackstone, private) | 28+ sites; large hyperscale wholesale |
| Iron Mountain Data Centers | ~20 sites; underground Boyers PA legacy + new builds |
| Compass Datacenters | Build-to-suit hyperscale (private) |
| Vantage Data Centers | Wholesale hyperscale + edge (DigitalBridge) |
| EdgeConneX | Edge / 2nd-tier-metro focus |
| GDS / Chindata / VNET | Chinese wholesale players |
| STACK Infrastructure | IPI Partners (now Blue Owl Digital) — global wholesale |
The colo industry distinguishes retail colo (cabinets, cages, often <250 kW per customer) from wholesale (full data halls, megawatts per customer, hyperscaler customers).
4.3 Enterprise on-prem
Private data centers owned by non-IT companies — banks, telcos, insurance, manufacturing, healthcare, government. Often legacy purpose-built rooms inside HQ buildings; in long-term retreat as workloads migrate to hyperscale or colo. Still substantial: ~50 % of total IT workload remained on-prem in 2024 (Gartner). Banks (latency-sensitive trading systems), defense (classification), and process-control (OT / SCADA) workloads remain reliably on-prem.
4.4 Edge
Distributed facilities placed close to end users to reduce latency or backhaul cost — towers, factories, retail, 5G MEC sites. Typical builds 20–500 kW, often containerized or modular prefab.
| Operator/category | Notes |
|---|---|
| EdgeConneX, Vapor IO, Compass Edge | Edge colocation pure-plays |
| Tower-co edge (Crown Castle, American Tower, SBA) | 5G MEC inside cell-tower compounds |
| Telco MEC (AT&T, Verizon, T-Mobile, Vodafone, BT, Telefónica) | Carrier-operated edge for low-latency 5G services |
| Industrial edge (factory-floor pods) | Cisco UCS-X, HPE Edgeline, Dell PowerEdge XR, NVIDIA EGX |
5. Power architecture
5.1 Utility feed
Modern hyperscale buildings sit at 115–230 kV transmission with a dedicated substation. A 100 MW data hall draws ~125 MVA at 0.95 power factor; even a moderate 30 MW colo needs a 38 MVA service which typically arrives at 34.5 kV from a utility substation 1–3 miles away. Redundancy options:
- Single feed — Tier I/II facilities, single point of failure.
- Two utility feeds from same substation — same outage events (lightning, switchgear fire) take both down.
- Two feeds from diverse substations on separate transmission rings — concurrent-maintainable; Tier III+.
- Two feeds plus behind-the-meter generation — gas turbines, fuel cells, or (rare) nuclear SMRs; full islandable.
For very large AI builds, hyperscalers are increasingly building behind-the-meter generation because the local utility cannot upgrade the transmission ring fast enough. Examples: xAI Colossus’s 14 × Solar Turbines Mercury and Titan units (~200 MW gas, behind permitted limits, became a regulatory flashpoint in mid-2025); Microsoft’s announced Three Mile Island restart (2028); Google’s Kairos Power SMR PPA (2030+).
5.2 Medium-voltage distribution
After the utility, the building’s MV grid distributes 13.8 kV or 33 kV to MV switchgear lineups (Eaton VCB, Schneider PIX, Siemens NXAir, ABB ZX2 / UniGear). Pad-mount or interior switchgear feeds unit substations with step-down transformers (typically 1500–3000 kVA dry-type or oil-filled K-13 rated for non-linear loads). Two-winding designs feed a pair of 480 V (US) or 415 V (EU) low-voltage main switchboards.
For 2N facilities, the entire MV → LV stack is duplicated as the A side and B side, with no electrical interconnection. Every load downstream that requires resilience has both an A-cord and a B-cord and an ATS or STS in front.
5.3 UPS architecture
The uninterruptible power supply bridges the gap between utility loss and generator pickup (usually 10–60 seconds) and conditions the power for IT loads.
Double-conversion (online) UPS is the dominant topology: AC → rectifier → DC bus + battery → inverter → AC. Inverter always carries the load; battery floats on the DC bus. Provides clean sinusoidal output regardless of input quality.
| Topology | Use |
|---|---|
| Double-conversion online | Most data centers; ~95 % efficient modern (Eaton 93PR, Schneider Galaxy VL, Vertiv Trinergy/PowerUPS, ABB DPA, Riello Multi Sentry) |
| Line-interactive | Small commercial, not data center |
| Static-switch bypass (ECO mode) | Saves ~2 % efficiency by carrying the load on bypass during clean utility, reverts to online in <2 ms on disturbance; some hyperscalers use full-time online for risk management |
| Distributed UPS (rack-mount) | Per-rack UPS (Schneider APC Smart-UPS, Eaton 9PX), more common for edge |
| Lithium-ion central UPS | Now ~75 % of new MW installations; LFP cells dominant since 2022 |
| VRLA central UPS | Sealed lead-acid 10-year batteries; legacy installs |
| Flywheel UPS | Active Power, Vycon — sub-30-second ride-through, generator-first architectures |
| Rotary UPS (Hitec, Piller) | Diesel-rotary UPS — engine, alternator, choke coil, and flywheel as one unit; eliminates batteries; ~3 % efficiency gain; common in EU |
The lithium-ion shift (2020 onwards):
VRLA (valve-regulated lead-acid) was the dominant UPS battery from the 1990s through 2018. Since 2020, ~75 % of new hyperscale UPS deployments use LFP (lithium iron phosphate) instead. Drivers:
- Footprint — Li-ion delivers 4–6× the energy density per cubic meter. A 1 MW Li-ion UPS battery cabinet replaces 4–6 VRLA cabinets, freeing whitespace.
- Cycle life — Li-ion: 5000–7000 cycles to 70 % capacity. VRLA: 500–1500 cycles. Real impact: VRLA needs replacement every 4–6 years (often before 10-year vendor claim); LFP makes 15-year designs realistic.
- Charge/discharge rate — Li-ion handles 5C+ rates so the same energy capacity also supplies much higher peak power.
- Thermal stability of LFP specifically — LFP cells avoid the thermal-runaway profile of NMC/NCA cells, which makes them code-acceptable inside the white space (NFPA 855-2023; UL 9540A; UL 1973).
- Temperature operating range — VRLA loses life sharply above 25 °C; LFP tolerates 35 °C. Frees the battery room from tight HVAC control.
Manufacturers: CATL, BYD, EVE, LG ES, Samsung SDI ship cells; UPS OEMs (Vertiv, Eaton, Schneider, ABB) integrate cabinet-level systems.
Battery room engineering:
- VRLA rooms — require ventilation per IEEE 1635 and 484 (hydrogen evolution under overcharge); EHS code requires <1 % H₂ in air with mechanical exhaust.
- Li-ion rooms — NFPA 855 sets thresholds (20 kWh per cabinet, 600 kWh per room without special protection); UL 9540A test data must be submitted for AHJ approval. Each cabinet has internal thermal-runaway detection; the room has linear heat detection + aspirating smoke (VESDA) + NOVEC 1230 or aerosol fire suppression; gas detection for off-gassing precursors (HF, CO).
- Spacing and access — IFC 1207 and NFPA 855 minimum 3 ft (0.9 m) aisle, 3 ft separation between cabinets, 100 sq ft (9.3 m²) maximum cabinet group without firewall.
5.4 Generator backup
When utility fails, the generator picks up the load. Standard architecture: each UPS is fed by a distribution panel which is in turn fed by an automatic transfer switch (ATS) that selects between utility transformer output and generator output. ATS transfer takes 5–30 seconds depending on engine ramp; the UPS battery bridges the gap.
Generators are almost always diesel (Caterpillar 3500/C-series, Cummins QSK, Kohler KD, MTU 20V4000/16V4000, Rolls-Royce/MTU). Newer “natural gas + diesel pilot” dual-fuel options exist for emissions-restricted urban sites. Sizes range from 1.5 MVA up to single 3.5 MVA paralleled banks; a 100 MW hyperscale has 30–40 generators in N+1 or 2N configurations.
EPA Tier 4 Final (40 CFR 1039) emissions standards apply to new stationary diesel >560 kW: requires SCR + DPF aftertreatment. Many AHJs grant 100 hour/year emergency exemptions that allow Tier 2 engines to remain in service. Local air permits (e.g., Virginia DEQ, Oregon DEQ) cap annual run-hours and require quarterly reporting.
Fuel storage — typically 24–72 hours at full load. A 100 MW facility burns ~10,000 gal/hour diesel; a 24-hour buffer = 240,000 gal in above-ground or underground tanks, secondary-contained per 40 CFR 112 SPCC.
5.5 Switching and downstream distribution
After generator + UPS, AC distribution goes to:
- Static Transfer Switches (STS) at the panel or PDU level for dual-corded loads. Subcycle transfer (<8 ms) between A and B feeds.
- Power Distribution Units (PDUs) — floor-mount transformers and panelboards (Vertiv Liebert PDU, Eaton ATC, Starline Track Busway). Output 415 V WYE (208 V to rack PDU) in EU/most-hyperscale, 480 V → 208 V step-down in US conventional.
- Busway / busbar — overhead aluminum busway (Starline, Anord Mardix, Universal Electric, Eaton, Schneider Canalis, Legrand DBO) is the dominant in-row distribution method for hyperscale: ratings 250–1200 A, snap-on tap-off boxes feed rack PDUs.
- Remote Power Panels (RPP) — fed from PDU, distributes branch circuits to row-end busway or directly to racks.
- Rack PDU — at the rack, 0U vertical (Server Technology, Raritan/Legrand, Schneider, Vertiv Geist). 30 A, 60 A, or 100 A inputs at 208/415 V; metered, switched, or basic. Modern intelligent PDUs report per-outlet kW and temperature/humidity to DCIM over SNMP/Modbus/Redfish.
5.6 OCP and high-voltage DC
The Open Compute Project standardized 12 V (later 48 V) DC distribution within the rack: a centralized rectifier shelf converts 208/415 V AC to 12 V or 48 V DC and feeds the servers directly, skipping per-server PSUs. Saves 3–5 % on conversion losses and removes per-server fans.
OCP V3 (Open Rack) standardized 21” wide racks at 48 V — adopted by Meta, Microsoft, and the OCP community. Hyperscale-internal AI racks (Microsoft’s “MX” series, Meta’s “Catalina/Grand Teton” for H100/H200, NVIDIA GB200 NVL72) all use 48 V DC bus internally.
400 V DC and 800 V DC are emerging for AI racks where 48 V bus current at 120 kW would exceed practical busbar copper. ABB, Eaton, and Vertiv ship 400 VDC PDU lines used in NVIDIA’s GB200 + GB300 reference designs.
6. Cooling architecture
6.1 The cooling problem
Every watt entering the building as IT power exits as heat. A 50 MW data hall must reject 50 MW of heat (170 million BTU/h, 14,200 tons of refrigeration) at a temperature lift acceptable to the silicon. Modern CPU/GPU junction temperature limits: ~95 °C for CPUs, 95–105 °C for GPUs. Working backward through the thermal stack (junction → die → package → heat sink or cold plate → coolant → outside air) sets the entering air or water temperature: typically 22–27 °C air or 35–45 °C water at the IT inlet.
ASHRAE TC 9.9’s Thermal Guidelines for Data Processing Environments (5th ed., 2021) define operating envelopes:
| Class | Recommended T (°C) | Allowable T (°C) | Use |
|---|---|---|---|
| A1 | 18–27 | 15–32 | Legacy enterprise mission-critical |
| A2 | 18–27 | 10–35 | Standard enterprise |
| A3 | 18–27 | 5–40 | Volume servers, modern cloud |
| A4 | 18–27 | 5–45 | Volume servers, max envelope |
| H1 | 18–22 | 18–25 | High-density (AI), recommended only |
| W17–W45 | 17–45 (water) | Up to 45 °C facility water | Liquid-cooled IT |
The W-class designations (water) introduced in the 2021 revision codify liquid-cooled operating points and have become the design reference for AI builds.
6.2 Legacy air cooling
CRAC (computer room air conditioner) — direct-expansion refrigerant compressors inside the unit. Liebert, Stulz, Vertiv. Mostly displaced in new builds.
CRAH (computer room air handler) — chilled-water coil only, refrigeration plant is centralized. Standard for enterprise and many colos; ~60–200 kW per unit.
Raised-floor underfloor supply — cold air pumped under 600–1200 mm raised floor, exits through perforated tiles in front of cabinets. Now considered legacy for densities above 8 kW/rack; airflow management becomes the constraint.
6.3 Hot-aisle / cold-aisle containment
Racks arranged in alternating rows with fronts facing each other (cold aisle) and backs facing each other (hot aisle). Containment physically separates the two with curtains, doors, and roof panels.
- Hot-aisle containment (HAC) — exhaust air captured at the rear of racks, returned to CRAH; the rest of the white space is cold supply temperature. Best for high-density.
- Cold-aisle containment (CAC) — supply air contained in front of racks; hot exhaust mixes in the white space and returns to the CRAH. Common retrofit.
A properly containment-sealed deployment can run a 25 °C supply air with ASHRAE A1 reliability and PUE 1.30 even on legacy equipment.
6.4 In-row and rear-door cooling
In-row CRAH (Schneider InRow, Vertiv CRV, Stulz CyberRow) — narrow cabinet sized to fit between racks, dramatically shortens the air path. Useful to 30–50 kW/rack.
Rear-door heat exchangers (RDHx) — passive (no fans) or active chilled-water coil bolted to the back of a rack. Removes 30–80 kW per rack by intercepting exhaust air before it leaves the rack. Coolflo/Motivair, ColdLogik, USystems, Vertiv Liebert XDR. Now the workhorse for 50–80 kW air-cooled AI racks (NVIDIA HGX H100 generation).
6.5 Direct-to-chip liquid cooling (DTC / D2C)
Cold plates clamped to CPU and GPU dies, supplied by a manifold and coolant distribution unit (CDU). Single-phase water-glycol working fluid at 35–45 °C supply, 45–55 °C return. Removes 70–100 % of the rack’s heat directly to liquid; the residual 0–30 % (RAM, NICs, PSU, switches) is handled by air-cooling.
The mandatory transition — NVIDIA’s GB200 NVL72 architecture (the 72-GPU rack-scale unit) is a 120 kW rack that ships liquid-only. Air cooling at this density is physically impossible (the boundary-layer ΔT exceeds the junction-to-ambient budget). 2024–2025 has been the year all hyperscalers gained operational liquid-cooling experience at fleet scale.
CDU vendors: Motivair (CoolDistribution series, Schneider acquired Sept 2024), Vertiv CoolPhase, CoolIT Systems CHx, Asetek InRack, DCX, JetCool, ZutaCore (two-phase), Iceotope.
Standards-emerging:
- OCP Open Liquid Cooling specifications (Universal Quick Disconnects UQD, ~31 manufacturers compliant); GB200 reference design uses Staubli SVC QDs.
- ASHRAE 90.4-2022 data-center energy code now includes liquid systems.
- TIA-942-C (2024) added Annex G covering liquid cooling.
6.6 Immersion cooling
The entire server is submerged in a dielectric fluid.
Single-phase immersion — synthetic hydrocarbon (Shell DC, Castrol DC, ExxonMobil Esso DC) or polyalphaolefin (PAO) coolant at 40–60 °C, circulated to an external dry cooler. Rack lays on its side as a tank (Submer SmartPod, GRC ICEraQ, Asperitas, LiquidStack DataTank, Iceotope KUL Edge, TMGcore EdgeBox). PUE 1.03–1.10 achievable. Heat capacity ~1.7 kJ/kg·K — much lower than water but acceptable with high mass flow.
Two-phase immersion — fluorocarbon working fluid (3M Novec 7100/7000 family, Chemours Opteon SF, Solvay Galden HT-55/110/135) boils at ~50–60 °C on the chip surface, condenses on a chilled coil above the bath. Excellent heat transfer (latent of vaporization), but 3M’s late-2022 announcement of PFAS production exit (Novec 7100 family scheduled to wind down by end-2025 in EU and major US sites under EPA PFAS restrictions) plus Chemours/Solvay PFAS legal exposure plus EU REACH proposals to restrict PFAS broadly have nearly frozen two-phase immersion deployments. The dominant survivor is single-phase + DTC hybrid.
6.7 Heat rejection — outside
Whatever the in-room cooling, the heat ultimately leaves the building through:
- Air-cooled chillers (Trane RTAF, Carrier 30RB, York YVAA) — vapor-compression to ambient air. Range: −20 °C to ~50 °C ambient. Roof-mount or pad-mount.
- Water-cooled chillers (Trane CenTraVac CVHF, Carrier 19DV, York YK, Daikin Magnitude) + cooling tower (BAC, Marley, Evapco, SPX, Mitsubishi). Tower evaporates water to ambient at the wet-bulb temperature, providing a lower condenser temperature than ambient air. Plant SEER 0.45 kW/ton vs 0.65 kW/ton air-cooled = ~30 % energy savings.
- Dry coolers / fluid coolers (BAC, Evapco, Frigel, Güntner) — closed-loop tube-and-fin heat exchanger; rejects water-glycol to dry air. No evaporative consumption but limited to ambient + 5 K approach.
- Adiabatic dry coolers — dry cooler with spray pre-cooling section for peak summer days. Hybrid balance of consumption and capacity.
- Free cooling / economizer modes — when outdoor wet-bulb is below CHW supply temperature, the chiller compressor is bypassed and the cooling tower provides cooling directly to the CHW loop. Climate-dependent: Hamina FI and Luleå SE run economizer 95 % of the year; Phoenix AZ <10 %.
6.8 Climate selection and site water
Free-cooling climate zones (sub-10 °C average wet-bulb 60 %+ of year) — Pacific Northwest US (Hillsboro, Quincy, Prineville, Umatilla), Iowa (Council Bluffs), Northern Virginia (winter only), Ireland, Netherlands, Nordic countries, Iceland. Sites are selected partially for these climate envelopes — Hamina FI and Luleå SE designs that pump seawater directly through CHW heat exchangers achieve PUE <1.10 year-round.
Water-stressed climates — Phoenix AZ, Las Vegas NV, Madrid, Singapore, Dubai. Public opposition to evaporative cooling has pushed builds to closed-loop dry or adiabatic. Microsoft’s 2024 commitment: zero potable water consumption for cooling by 2030; Google’s: water-positive by 2030; Meta’s: 100 % new builds dry-cooled in water-stressed regions.
7. Networking architecture
7.1 Spine-leaf
The 2010s-vintage Clos topology — every leaf (top-of-rack) switch connects to every spine switch — is the dominant data-center fabric. Replaces three-tier access/aggregation/core with two layers; provides equal-cost-multi-path (ECMP) load balancing for east-west traffic.
A modern hyperscale “data hall” pod:
- 32 racks × 1 ToR switch = 32 leaf switches
- 8 spine switches above
- 16 × 400G uplinks per leaf to spines = 6.4 Tbps leaf uplink
- Servers: 48 × 25/50/100 GbE per leaf
Multiple pods aggregate into a “super-spine” for inter-pod traffic. Hyperscale fabrics scale to 100k+ servers in one administrative domain.
7.2 Underlay routing — BGP
Modern hyperscale fabrics run eBGP-over-pod as the underlay (RFC 7938: “Use of BGP for Routing in Large-Scale Data Centers”). Each ToR is its own ASN; each spine is its own ASN; routes are advertised host-route-precise. Replaces OSPF/IS-IS for convergence-time and policy reasons. Implementations: FRR (Free Range Routing), Cumulus / NVIDIA Linux Switch (later renamed NVIDIA Air), SONiC.
7.3 EVPN-VXLAN overlay
Network virtualization for multi-tenant + multi-VRF — EVPN (Ethernet VPN) signaled via MP-BGP, VXLAN (Virtual Extensible LAN) for L2-over-L3 tunneling. Encapsulates tenant VLANs / segments in 24-bit VNI (16M segments) and tunnels them on top of the BGP-routed underlay. Standard for service-provider edge and cloud-multi-tenant.
7.4 Cabling — copper, fiber, optics
| Medium | Use |
|---|---|
| DAC (direct-attach copper) | <3 m server-to-ToR or ToR-to-spine; passive 25–400G; lowest cost, lowest power |
| AEC (active electrical cable) | 3–7 m extension; built-in retimer; lower power than optics for short reach |
| ACC (active copper) | Replaced largely by AEC; some 400G variants |
| Multi-mode fiber (OM4/OM5) | Inside building, 100–400G QSFP28/QSFP-DD, OM5 supports SWDM4 to 150 m |
| Single-mode fiber (OS2) | Long-reach inside-building + campus + WAN |
| Active optical cable (AOC) | Permanently-attached optics; common 100/400G |
| 400G/800G optics | 400G-FR4, 400G-DR4, 400G-DR1 single-lambda 100G+; 800G-DR8/FR8 in 2024 deployments; 1.6T standardized for 2026 |
| CPO (co-packaged optics) | Optics integrated with switch ASIC die; reduces SerDes power, in early shipments 2024–2025 |
| LPO (linear pluggable optics) | Optics without DSP; halves optic power; 51.2 Tbps switches use LPO 800G |
ToR switches in 2024–2025 are 51.2 Tbps (64 × 800G or 128 × 400G), built on Broadcom Tomahawk 5, NVIDIA Spectrum-4, or Cisco Silicon One G200. The next generation (102.4 Tbps, late 2025/2026) is built on Tomahawk 6 and Spectrum-5.
7.5 NIC fabrics — RoCE and InfiniBand for AI
AI training clusters use a separate “back-end” fabric dedicated to GPU-to-GPU all-reduce traffic. Two technologies dominate:
- InfiniBand — NVIDIA Quantum-2 (NDR 400 Gbps) and Quantum-X800 (XDR 800 Gbps). Hardware credit-based flow control gives lossless behavior. NVIDIA’s preferred fabric (the Mellanox acquisition cemented this); used in DGX SuperPOD reference designs and the largest 2024 AI builds.
- RoCEv2 (RDMA over Converged Ethernet) — RDMA-capable Ethernet using PFC (priority flow control) and ECN for lossless. Ethernet-switch agnostic (Broadcom, NVIDIA Spectrum, Cisco Silicon One, Marvell Innovium, NVIDIA Spectrum-X for AI-specific congestion control). Microsoft, Meta, Google, AWS run RoCEv2 fabrics at multi-100k-GPU scale.
The Ultra Ethernet Consortium (UEC, July 2023) is standardizing AI-targeted Ethernet evolutions; first compliant gear shipping 2025–2026.
7.6 Cabling infrastructure standards
TIA-942-C-2024 specifies cabling, pathways, redundancy classes (Rated 1–4 mirroring Uptime tiers), and physical security. BICSI 002-2024 Data Center Design and Implementation Best Practices is the planning companion. ANSI/TIA-568.3-E covers fiber categories (OM3/4/5, OS1/2) and termination. ANSI/TIA-606-C addresses identification, labeling, and administration.
8. Tier classification — Uptime Institute
The Uptime Institute Tier Standard is the dominant facility classification, certifying Design + Constructed Facility + Operational Sustainability:
| Tier | Description | Single-failure tolerance | Concurrent maintenance | Availability target |
|---|---|---|---|---|
| Tier I | Basic site infrastructure | No — any failure or maintenance drops IT | No | 99.671 % (28.8 h/yr) |
| Tier II | Redundant capacity components | Power/cooling can tolerate one component failure if redundant unit picks up | No — maintenance requires shutdown | 99.741 % (22.0 h/yr) |
| Tier III | Concurrently maintainable | All distribution paths and equipment have alternates; any single component or path can be taken out of service for maintenance without dropping IT | Yes | 99.982 % (1.6 h/yr) |
| Tier IV | Fault tolerant | Active/active distribution paths; any single failure event leaves no impact on IT; includes compartmentalization to limit fire/leak/water-event spread | Yes | 99.995 % (26 min/yr) |
A facility achieves Tier III by being “concurrently maintainable” — there are A and B distribution paths and one can be taken offline for service. Tier IV adds “fault tolerance” — both paths can be hot, both protected, both compartmentalized. Tier IV is rare (~80 sites globally certified) and the marginal cost over Tier III rarely pencils outside financial-services trading and certain healthcare.
TIA-942-C “Rated 1–4” is a parallel scheme with similar levels; the two coexist commercially.
EU has EN 50600-1/-2/-3/-4 which defines availability classes 1–4 with similar intent.
9. Site selection
Site selection is a multi-criteria optimization. Practical hierarchy:
- Power availability — can the local utility deliver 100+ MW in 24 months? Most denied; Loudoun, Northern Virginia took a 12+ year pipeline by 2024 because of this. Hyperscalers now look for sites with existing transmission rights, retired generation interconnect rights, or behind-the-meter generation capability.
- Power cost — at 50–60 % of OPEX, 1 ¢/kWh × 100 MW × 8760 h = \$8.76M/year. Hyperscaler PPAs target 3–5 ¢/kWh; renewables PPAs lock in 20+ years. Iowa wind, Nordic hydro, Pacific Northwest hydro all materially cheaper than ERCOT, PJM, or CAISO.
- Carbon profile of grid — hyperscalers’ 100 % CFE / 24-7 matching commitments add 1.5–2x weighting to power-mix decisions.
- Water availability (if evaporative) — increasingly a binding political constraint in arid US Southwest.
- Fiber routes — diverse path count, distance to peering points. Northern Virginia’s dominance is partly that 80 % of US East Coast fiber landings converge there.
- Latency proximity to user population — drives edge and “tier-2 metro” build (Columbus, Atlanta, Reno, Salt Lake).
- Climate — wet-bulb temperature for free cooling and water consumption; storm and seismic hazard exposure.
- Land — 100+ acre flat parcels with road access; campus master plans of 300–1000 acres.
- Hazards — flood zone (FEMA Zone X preferred; never AE/VE), seismic (UBC 2, IBC SDC A/B preferred), wildfire (WUI exclusion), tornado, hurricane, ash zones near volcanos, geological subsidence.
- Tax incentives — most hyperscaler builds rely on a sales-tax exemption on IT equipment (Virginia, Texas, Iowa, Ohio, Oregon, Washington) and property-tax abatement (10–30 years).
- Workforce — operators (technicians, electricians, engineers) within commute; community college pipelines.
- Permitting and regulatory — local zoning, state environmental review (NEPA-equivalents), utility interconnection studies that may take 1–3 years independent of build.
- Security and geopolitics — sovereign-residency rules (EU GDPR, China cybersecurity law, India data-localization), proximity to defense installations, security clearance availability.
10. AI-specific design (2024–2026)
The 2023 LLM scaling cliff fundamentally changed data-center design. The traditional 5–15 kW rack with air cooling is incompatible with NVIDIA H100/H200/B200/GB200/GB300 reference designs.
| Generation | Per-GPU TDP | Common rack TDP | Cooling |
|---|---|---|---|
| H100 SXM (Hopper, 2022) | 700 W | 30–50 kW (HGX × 4–8 servers) | Air OK, RDHx common |
| H200 SXM (2024) | 700 W | 30–50 kW | Air OK |
| B200 SXM (Blackwell, 2024) | 1000 W | 60–80 kW | Air or DTC |
| GB200 NVL72 (Blackwell, 2024–2025) | 1200 W | 120 kW per NVL72 rack | DTC mandatory (>72 GPUs liquid-coupled with NVLink switches) |
| GB300 NVL72 (Ultra, 2025–2026) | 1400 W | 140 kW | DTC + facility water at 35–45 °C |
| Rubin / R200 (planned 2026–2027) | 1800 W estimated | 200+ kW | Likely facility water + cold-plate optimized |
| AMD MI300X (Instinct, 2024) | 750 W | 40–60 kW | Air/RDHx |
| AMD MI350X (2025) | 1000 W estimated | 60–80 kW | DTC |
Network: AI clusters require a second fabric just for GPU-to-GPU traffic. A 32k-GPU H100 cluster has ~32k front-end NICs (Ethernet for storage/management) plus ~256k back-end ports (8 InfiniBand NICs per server). The back-end network alone can consume more port-count than the front-end.
Storage: AI training I/O patterns are read-mostly large-block (checkpoints) plus shuffled small-block (dataset). All-flash NVMe-oF (NVMe over Fabrics) on Ethernet/IB is dominant; legacy SAN architectures fall over. DDN, VAST Data, Weka.IO, IBM Storage Scale (Spectrum Scale / GPFS), Pure Storage FlashBlade dominate. Capacities: 50–500 PB usable.
Power profile: AI training has a much higher utilization (80–95 % continuous) than traditional cloud (30–50 %). A 100 MW AI hall actually consumes ~85 MW continuous vs ~40 MW for an equivalent cloud hall. UPS sizing, generator sizing, and PPA load forecasting must reflect this.
11p. Edge cases / gotchas
- Resource adequacy and ride-through — Generators take 5–30 s to come online; UPS battery autonomy is sized to that plus generator-test margin. A failed generator start at the wrong time = full UPS depletion = blacked rack. Modern best practice: cross-tie generators across multiple buses with paralleling switchgear, plus seasonal load-bank testing.
- Black-start dependencies — Some generator controls require house AC for fuel pumps, jacket-water heaters, starting batteries trickle charging. If utility fails during a generator maintenance cycle, the generator may not start. Cross-redundancy of these “house aux” loads is a Tier IV requirement.
- Single point of failure cabinets — A common audit finding: PDU and UPS cabinets share a single distribution panel between them, defeating the 2N architecture. Diverse routing of A and B cables through separately fire-rated conduits is essential.
- Battery thermal runaway — One LFP cell event can cascade if the cabinet’s internal HVAC fails. 2022 Iberdrola APS Surprise AZ event (lead-acid) and the 2024 Korean BESS fires (NMC) remind designers that thermal monitoring + early-detection venting must be fail-safe.
- Cooling-water leaks in liquid-cooled IT — A single coupler failure can drip onto a $300k GPU board. Leak-detection cabling (TraceTek, RLE), drip pans, and dielectric-fluid options reduce risk; insurance contract terms now distinguish “wet” vs “dry” facilities.
- PFAS phaseout — Two-phase fluorocarbon immersion fluids (Novec/Galden/Opteon SF) face shrinking availability; design teams that locked in two-phase architectures in 2022 are now refitting for single-phase + DTC.
- Concurrent maintainability is paperwork — Tier III certification depends on demonstrating that any component can be removed without impact. Field testing shows 5–10 % of newly built Tier III sites fail a concurrent-shutdown test on first audit. Detailed switching procedures + load-shed protocols + operator training are as important as topology.
- Generator emissions permits — 100 hours/year of testing + emergency operation is common; if a year exceeds that (an extended outage, repeated brownouts), the facility may be in violation of its air-permit. Some 2022–2023 Texas heat dome events pushed several DCs into permit-exceedance territory.
- Harmonics and neutral current — IT loads are non-linear; THD_i 25–40 % is common without mitigation. Neutral conductor sized 200 % of phase or use of 18-pulse or active filters per IEEE 519-2022 mandatory. Triplen harmonics (3rd, 9th, 15th) add on the neutral.
- DCIM and BMS gaps — BMS (building management) and DCIM (data-center infrastructure management) often interoperate poorly. BACnet/IP, Modbus TCP, SNMP, MQTT, Redfish, Niagara — a 30 MW facility may have 8+ protocol stacks needing integration. SiteWise, Sunbird, Nlyte, Schneider EcoStruxure IT, Vertiv Trellis are the dominant DCIM platforms.
- Stranded capacity — A common operational pathology: power, cooling, and rack space sized correctly for design, but actual load profiles diverge (cooling at 60 % utilization, power at 90 %, space at 110 % — one ahead, one behind, one out of room). Reserving a fungible buffer across all three is a planning best-practice but rare in execution.
- Floor loading — A loaded Open Compute Rack V3 with 48 V batteries can hit 1200–1500 kg per cabinet. Slab loadings ≥1500 kg/m² (≥300 psf) live load are now standard for AI builds. Older retrofit data halls (≤500 kg/m²) often need slab reinforcement before liquid-cooling deployment.
- Sustainability accounting — The PUE/WUE/CUE/REF scorecard is being formalized into mandatory disclosures: EU’s Energy Efficiency Directive (EED) 2023 recast requires DCs >500 kW to report quarterly; SEC climate disclosure rule (final March 2024, litigation pending); California SB 253/261. Voluntary frameworks: SBTi, GHG Protocol Scope 2 location-based vs market-based.
- Permitting and community pushback — Loudoun County’s 2023 zoning moratorium, Memphis’s xAI emissions complaints (2025), Phoenix water disputes — community opposition is the new top-3 schedule risk for hyperscale builds, ahead of supply chain and labor.
12p. Tools & software
Design / analysis — AutoCAD MEP, Revit MEP, Bentley OpenBuildings, IES Virtual Environment, IES VE-DataCenter, Future Facilities 6SigmaDCX (CFD), Cadence Reality Digital Twin (CFD, ex-Future Facilities), Romonet, NREL Data Center Energy Profiler, ASHRAE DOE-2 / EnergyPlus for shell.
Power-systems analysis — ETAP, SKM PowerTools, EasyPower, Power Factory (DIgSILENT), Aspen OneLiner.
Cooling CFD — Cadence 6SigmaDCX (the dominant DC-specific CFD), Ansys Fluent, Siemens STAR-CCM+ (with DC-specific add-ons).
BMS / building automation — Tridium Niagara, Siemens Desigo CC, Honeywell WEBs / JACE, Schneider EcoStruxure Building Operation, Johnson Controls Metasys.
DCIM / asset management — Sunbird dcTrack, Nlyte, Schneider EcoStruxure IT Expert, Vertiv Trellis / Environet Alert, Cormant CS, Device42, NetBox + Nautobot (open-source), Hyperview, FNT Command, Cisco UCS Director.
Standards / commissioning — Uptime Institute Tier Standard, ASHRAE TC 9.9 Guidelines, TIA-942-C, BICSI 002, EN 50600, ISO/IEC 22237.
Sustainability / monitoring — Schneider Resource Advisor, Microsoft Cloud for Sustainability, Persefoni, Watershed.
13. Cross-references
- transformers-power-systems — utility tie, station-class transformers, generation interconnection.
- electric-motors — chiller compressors, pump motors, fan motors; IE3-IE5 efficiency, VFD compatibility.
- power-electronics — UPS topologies, VFD harmonics, rectifier/inverter architectures.
- hvac-fundamentals — chiller plant, cooling tower, psychrometrics, refrigerants, IAQ.
- heat-transfer — coil UA, cold-plate analysis, immersion convection, junction-to-ambient resistance budgets.
- refrigeration-cycles — chiller vapor-compression detail, COP vs PUE.
- fluid-mechanics — duct + pipe pressure drop, pump head, manifold design for liquid-cooled racks.
- pumps-turbomachinery — chilled-water and condenser-water pumps, cooling-tower fans.
- digital-control — BAS PID loops, supervisory sequencing, chiller plant optimization.
- reliability-engineering — MTBF / MTTR, redundancy classes, FMEA, Markov availability modeling.
- cybersecurity-engineering — OT/ICS security for BAS, DCIM, physical access.
- standards-bodies — Uptime Institute, TIA, BICSI, IEEE, ASHRAE, ISO/IEC.
- battery-chemistries — LFP, NMC, NCA, VRLA chemistry detail.
- energy-storage-systems — BESS sizing, augmentation cycles, lifetime.
- refrigerants — A2L transition, GWP, phase-out timeline.
- heat-transfer-correlations — internal forced convection, cold-plate Nusselt correlations, two-phase boiling.