Originally published by DataCentre Magazine

Thomas Steen, CTO of Nortek Data Center Cooling, explains why the data centre industry needs a shift in thermal management strategy

Traditional air-cooling systems approach their limits as AI workloads drive unprecedented increases in rack power and thermal density. 

Today’s server racks operate in the 50kW to 250kW range. However, next-generation deployments are expected to approach the megawatt scale soon. 

As compute density accelerates, conventional air-cooling architectures can no longer deliver the thermal performance, efficiency or scalability required to support future infrastructure demands. 

In this byline for Data Centre Magazine, Thomas Steen, CTO of Nortek Data Center Cooling (Nortek DCC), argues that the industry needs a fundamental shift in thermal management strategy.

Thomas Steen, CTO of Nortek Data Center Cooling

Understanding the challenge

Training large language models (LLMs)  may require thousands of GPUs operating continuously for months at a time, resulting in unprecedented levels of energy consumption. GPUs designed for AI training and inference can consume more than ten times the power of conventional processors, dramatically increasing rack-level power density and the associated thermal load. 

As rack densities continue to rise, thermal management becomes mission-critical to system performance, hardware reliability, and overall facility uptime. Inadequate cooling introduces operational risk, including thermal throttling, accelerated hardware degradation, unexpected downtime, and potential failure of multi-million-dollar infrastructure.

The brisk adoption of AI has also sped up the deployment of high-performance accelerated servers, increasing the power density in data centres, which we have observed in many deployments. Greater power density raises temperatures. Air cooling is not enough.

Innovative thermal management is crucial for data centre efficiency as AI demand rises (Credit: Nortek Data Center Cooling)

Air cooling’s role isn’t disappearing. It’s evolving. In next-generation AI and HPC environments, air cooling will serve as a supplemental technology, supporting ancillary room loads and non-liquid-cooled components, while remaining the dominant cooling method for traditional enterprise data centres. However, air cooling alone is no longer capable of supporting the thermal demands of modern AI workloads at scale.

As rack power densities exceed the 40 kW to 50 kW threshold, maintaining uniform airflow distribution becomes increasingly difficult, resulting in recirculation, thermal stratification, and localised hotspots that can compromise system reliability and increase the risk of hardware failure. 

When heat cannot be controlled

Heat management is a delicate balance. Every material in a data centre has a distinctive coefficient of thermal expansion (CTE). Semiconductor substrates and copper have very different CTEs, for example. Data centre designers and operators must consider each material’s unique responses to rising temperatures. This is especially challenging with AI racks, since AI workloads not only generate massive amounts of heat but are also unpredictable.

Liquid cooling is the only scalable solution capable of supporting next-generation AI infrastructure.  Liquid cooling is the only scalable solution capable of supporting next-generation AI infrastructure. 

Thomas Steen, CTO of Nortek Data Center Cooling

If CTE differences are not addressed, soaring temperatures can expose design weaknesses and cooling limitations, which may cause mechanical problems, including cracked solder joints and board warping. Elevated temperatures impact device reliability. If the internal temperature of a chip exceeds its maximum rating, its performance suffers. 

Mechanical failures can knock out entire systems and crash services. The industry saw a large-scale example of this consequence a few years ago when a hyperscaler outage knocked down a cloud service across Western Europe. The cause of the outage was a “thermal event”, which means the cooling systems failed. 

Liquid cooling is necessary for AI factories

In “AI factories,” liquid cooling is a critical infrastructure requirement rather than an optional technology. Compared to traditional air cooling, liquid cooling offers superior thermal transfer capabilities and enables heat removal directly at the source. This direct-to-source approach is essential for managing the ballooning heat densities.

Team members at the Dyersburg facility (Credit: Nortek Data Center Cooling)

Because liquids are superior heat conductors to air, liquid cooling systems operate with greater efficiency, reducing cooling energy consumption, lowering operating costs and improving thermal performance. In addition, liquid cooling helps reduce water usage and support sustainability objectives, which is increasingly important as data centres face growing resource constraints. 

The transition to liquid cooling

The transition to liquid cooling is certain. Liquid cooling is the only scalable solution capable of supporting next-generation AI infrastructure. 

However, air cooling will continue to play an important supporting role through hybrid cooling architectures that combine both technologies. Solutions such as a combination of rear door heat exchangers, dry coolers and retrofit liquid cooling systems can help existing facilities support some level of higher-density deployments without requiring a complete overhaul.

Most new AI-focused data centres will require liquid cooling by design, while many legacy facilities may need to be retrofitted as workload demands evolve. Traditional air cooling alone is no longer sufficient for the thermal demands of AI and HPC environments, reinforcing the industry-wide shift toward liquid cooling strategies.