HomeSecurityNvidia Blackwell at Microsoft: Cooling Called “Wasteful”

Nvidia Blackwell at Microsoft: Cooling Called “Wasteful”

Nvidia’s push to roll out its latest Blackwell hardware inside hyperscaler data centers is running into a familiar reality: the chips are fast, power-hungry, and they force operators to make tough trade-offs around cooling.

In an internal email sent in early fall, an Nvidia employee involved with a Blackwell installation at a Microsoft facility described the company’s cooling approach as “wasteful”—even while acknowledging it adds flexibility and fault tolerance, according to a memo viewed by Business Insider.

The note offers a rare peek into what it looks like when next-gen AI infrastructure meets the practical constraints of existing data center designs.

An NVIS email flags Microsoft’s cooling as “wasteful”

The email came from a staffer on Nvidia’s Infrastructure Specialists team and discussed the setup of two GB200 NVL72 racks—rack-scale systems built around 72 Nvidia GPUs each. The installation was described as supporting OpenAI workloads, with Microsoft serving as a key infrastructure partner for OpenAI.

Because densely packed GPUs generate enormous heat, the systems rely on liquid cooling at the server/rack level.

But the Nvidia staffer suggested Microsoft’s overall facility approach looked inefficient given the system’s size and the apparent lack of facility water use—while still noting it brings resilience.

Two layers of cooling, two sets of trade-offs

One way to interpret the comment is that it’s not just about what happens inside the racks.

Shaolei Ren, an associate professor of electrical and computer engineering who studies data center resource use, described a two-part setup: liquid cooling to move heat off the chips, plus a building-level system to expel that heat from the facility.

Ren suggested the “wasteful” remark could point to a building-level design that leans more heavily on air cooling rather than water-based heat rejection. The upside: less water use. The downside: air cooling can require more energy.

In other words, the industry’s cooling debate isn’t only about performance—it’s also about optics, local constraints, and which resource is scarcer where the building sits.

Microsoft: closed-loop liquid cooling added to air-cooled sites

Microsoft told Business Insider it uses a closed-loop liquid cooling heat exchanger unit that can be deployed in existing air-cooled data centers to increase cooling capacity for both Microsoft and third-party platforms.

The company framed the approach as a way to scale AI infrastructure quickly by maximizing its global footprint, while improving heat dissipation and power delivery for AI and hyperscale systems.

Why the “2× faster than Hopper” line needs context

Nvidia unveiled Blackwell in March 2024, positioning it as a major step up from the Hopper generation.

But performance claims around AI hardware are heavily workload-dependent—and comparisons can vary based on whether you’re talking training, inference, precision formats, power envelopes, and system configuration. So while “roughly twice as powerful” captures the direction of Nvidia’s messaging, it’s not a universal rule across every scenario.

Meanwhile, Nvidia is already talking up newer Blackwell configurations, including the GB300 generation, which is now being marketed as available.

Inside the installation: validation headaches and improving hardware quality

The Nvidia email also described the kind of friction that’s common when new hardware hits real data centers.

The staffer said onsite support was necessary, with significant effort going into validation documentation and making sure the steps were clear to teams less familiar with cluster and system validation. The handoff processes between Nvidia and Microsoft also reportedly needed more refinement than usual.

Still, the memo struck an optimistic note about the production hardware. The staffer wrote that GB200 NVL72 production units looked better than early samples, and both racks achieved a 100% pass rate on certain compute performance tests.

Nvidia, for its part, said its Blackwell systems deliver strong performance, reliability, and energy efficiency across a wide range of use cases. The company also said customers—including Microsoft—have deployed Blackwell at significant scale, though it’s more precise to discuss scale in terms of GPUs and deployed capacity rather than implying “hundreds of thousands” of full rack-scale NVL72 systems.

The bigger picture: cooling is becoming the next flashpoint

As AI buildouts accelerate, cooling choices are increasingly colliding with public scrutiny around energy and water consumption.

Ren described the decision as a resource trade-off: air cooling can reduce visible water usage (often a public pressure point), but it can raise energy demand—forcing operators to weigh cost, availability, and perception.

Microsoft has said it aims to become carbon negative, water positive, and zero waste by 2030, and has also discussed a zero-water cooling design for next-generation data centers. Whether those designs can scale cleanly alongside rapidly rising AI compute demand is one of the biggest infrastructure questions facing the industry right now.

RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

- Advertisment -

Most Popular

POPULAR TAGS

- Advertisment -