Glossary of terms used
- CDU
- Coolant Distribution Unit. Thermal system component distributing chilled coolant to racks or direct-to-chip cold plates.
- GPU
- Graphics Processing Unit. The compute silicon at the centre of AI workloads.
- SiC
- Silicon Carbide. Wide-bandgap semiconductor material used in high-voltage power electronics.
For the full corpus glossary of acronyms used across all essays, see adikumar.co/glossary.
The Thermal Stack: Technical Companion
Where the market essay named the vendors, this piece explains the physics. What 1,000 W/cm² means at the die, how a coupling that costs a hundred dollars either saves a rack or destroys it, why the industry is moving cooling from the room into the package, and where the technology goes next. Cross-sections, flow diagrams, and roadmaps for readers who want to understand the plumbing under the market.
Heat is not information
Every watt of electrical power a GPU consumes leaves the die as heat. There is no version of a switching transistor in which that stops being true, and no amount of software cleverness changes it. What determines whether the die throttles is how fast that heat gets to a fluid cool enough to accept it. That is a fluid mechanics and materials-science problem, and every layer in the thermal stack is a specific answer to it.
Three mechanisms move heat: conduction (through solids and stagnant fluids, described by Fourier's law q = -k·∇T), convection (through moving fluids, described by Newton's law of cooling q = h·A·ΔT), and radiation (from any surface above absolute zero, but negligible below 150°C). Data centre cooling is a game played almost entirely in the first two. Conduction dominates from the die to the coolant. Convection dominates from the coolant to the atmosphere.
Junction temperature is the master variable
Silicon devices are specified to a maximum junction temperature (T_j_max), typically 105°C for hyperscale AI accelerators, 125°C for automotive silicon, 200°C+ for SiCAbove T_j_max, one of three things happens: the device throttles (reduces clock frequency or blocks current), it drifts electrically (threshold voltage shifts, leakage rises), or it fails outright. The whole thermal design job is to keep T_j below T_j_max while the die dissipates its rated power.
The relationship is a simple series-resistance model. Heat flows from junction through a stack of thermal resistances (die, interface material, lid, second interface, cold plate, coolant film) into the coolant. Each layer adds a temperature drop equal to the heat flow times its thermal resistance. Add them up and you get the required coolant temperature to keep the junction cool enough.
CDU architecture, coolant chemistry and facility loop separation determine how efficiently heat leaves the chip package. Single-phase glycol handles up to 400W per die comfortably. Two-phase dielectric extends past 1,500W. Fluid degradation, secondary-loop maintenance and heat-exchanger fouling add operating cost that offsets some of the density gain.