Tech Week Singapore 2026
The Cooling Tax: How Agentic AI Turns Megawatts into More Tokens
AI factories are becoming power constrained, not compute constrained. Grid interconnects cap how much power a site can draw, so the only real lever left is how that fixed budget is split, and cooling can eat up to 35% of it today.
We built a chip to chiller data center simulation to test one thesis: every megawatt saved on cooling can be redirected straight into compute. Modeling a real 3360 GPU, roughly 2 MW NVIDIA B200 environment, we deployed eight zero shot LLM agents, six open weight and two proprietary, to make live cooling decisions, scored purely on tokens generated per joule. Every model beat a static baseline, with reward gains up to 11.5% and cooling overhead cuts up to 10.3%.
But efficiency numbers alone hid a real risk. As we gave agents more to control, more of them crossed a hard 85°C junction temperature safety limit, up to six of the eight models in our toughest test, even as their reward and efficiency scores looked just as strong as the safe ones. Tokens per joule cannot be the only number that matters.
This talk shares the architecture, the results, and what it will take to scale this from one 2 MW rack to gigawatt scale deployments, safely.
Cloud & AI Infrastructure
Cyber Security World
Big Data & AI World
Data Centre World 



























