Glossary of terms used
- CDU
- Coolant Distribution Unit. Thermal system component distributing chilled coolant to racks or direct-to-chip cold plates.
- EBITDA
- Earnings Before Interest, Tax, Depreciation, and Amortisation. The most commonly referenced operating-earnings metric in M&A pricing.
- GPU
- Graphics Processing Unit. The compute silicon at the centre of AI workloads.
- GaN
- Gallium Nitride. Wide-bandgap semiconductor material used in RF and power applications.
- SiC
- Silicon Carbide. Wide-bandgap semiconductor material used in high-voltage power electronics.
- TAM
- Total Addressable Market. The maximum revenue opportunity available if a product served every potential customer segment.
For the full corpus glossary of acronyms used across all essays, see adikumar.co/glossary.
The Thermal Stack: How Heat Became the Binding Constraint on AI Compute
Every watt delivered to a GPU comes back out as heat, and at 600kW per rack there is nowhere for it to go but water. A layer-by-layer analysis from the die to the dry cooler: the $9.5bn acquisitions, the coupling oligopoly that most models miss, the fluid chemistry stranded by a single corporate decision, and why the most defensible margins in the stack sit in components costing a hundred dollars.
- Liquid is no longer optional. NVIDIA's VR200 compute and switch trays are fanless; rack airflow requirements drop ~80% while coolant flow roughly doubles versus GB300. Microsoft has confirmed all future Maia deployments are liquid-default. The air-cooled hyperscale training cluster is over.
- The market is ~$5.5-6.8B in 2026 growing 18-26% depending on scope, inside a broader ~$18.5B data center cooling market. Cooling content per rack is rising ~12% per generation while power content rises ~32%.
- The chokepoint is a coupling. A GB200 rack uses 100+ universal quick disconnects. Stäubli, CPC and Parker held 80%+ of the Chinese UQD market as recently as 2024, and fewer than fifteen firms worldwide mass-produce complete units. Western UQDs sell at RMB 80-120 against RMB 30-50 domestic. That price umbrella is now under direct attack.
- PFAS regulation stranded two-phase immersion. 3M's PFAS exit ended Novec and Fluorinert production in 2025, collapsing the supply chain for two-phase. Single-phase held ~81% of immersion in 2024 and is gaining. The ECHA restriction opinion lands end-2026.
- Consolidation has been aggressive and expensive. Eaton / Boyd Thermal at $9.5bn (22.5x EBITDA), Schneider / Motivair (~$850M), Vertiv / PurgeRite (~$1bn). Cooling assets are clearing at multiples that assume the AI buildout does not stop.
- Profit pools cluster at the extremes. Couplings, TIMs and fluids (the cheapest components) hold the best margins because they are qualification-gated and failure-critical. CDUs and cold plates are where volume lives and where Taiwanese and Chinese competition compresses hardest.
- The frontier is moving inside the package. Microchannel lids, then microchannels etched into silicon itself, then backside liquid cooling. TSMC is integrating microchannel cooling into its 3DFabric platform. The cold plate's job is migrating onto the die.
- NVIDIA is tightening its grip on the thermal supply chain, standardising designs and squeezing supplier margins. That is the dominant structural risk to every independent vendor in this map.
- Bottom-up TAM. Roughly $5.5-6.5B in 2026 rising to $22-30B by 2030 in the base case, driven more by attach rate and content escalation than by facility count.
The series walks a single physical path. It begins at the medium-voltage utility bus at the site fence, steps down through the substation and switchgear, arrives at the datacenter rack where 800V DC is stabilised by the capacitor stack, is converted by silicon-carbide switches to 48V, is distributed across the rack by copper busbars and whips, is stepped down again by multi-phase controllers on the accelerator board to 0.8V, and finally routed through the on-package power delivery network to a transistor gate drawing over 2,000 amperesWaste heat from every conversion stage is removed by the thermal stack. The whole thing is packaged inside a factory-modular building because there aren't enough electricians to build it stick-frame. Six essays. One 800V → 0.8V staircase.
- Part I. The Capacitor Stack 800VDC at the rack
- Part II. The Wide-Bandgap Stack SiC and GaN conversion
- Part III. The Thermal Stack Removing the waste heat (you are here)
- Part IV. The Interconnect Stack Busbars and whips
- Part V. The On-Package Delivery Stack 48V to 0.8V
- Part VI. The Modular Datacenter Stack How the building gets built
The thermal value chain, layer by layer
The power chain divides by voltage. The wide-bandgap chain divides by process step. The thermal chain divides by position along the heat pathThe economics invert as you travel it: the components nearest the die are the smallest, cheapest and most defensible; the components nearest the atmosphere are the largest, most expensive, and most commoditised. Read the map with that in mind.
Above 100 kW per rack, direct-to-chip liquid cooling is the only viable AI data centre thermal architecture. The thermal stack has four layers: cold plate / immersion at the chip, coolant distribution units (CDU) for loop separation, facility heat rejection, and chemistry supply chain. CoolIT leads DLC by design-win share; adjacent M&A (Ecolab acquiring CoolIT, Eaton acquiring Boyd) is reshaping vendor concentration.