For two decades, data center operators built their electrical monitoring strategies around a simple assumption: power demand was predictable, distributed, and relatively stable. Rows of compute servers drew roughly equal power. Load grew gradually. Thermal patterns were well-understood.
AI has shattered every one of those assumptions.
A single rack of H100 GPUs can draw 80 kilowatts or more. A dense AI training cluster can pull megawatts from a single switchgear section. Load swings from near-zero to full capacity within minutes as training jobs start and stop. And the electrical infrastructure in most data centers was never designed for any of this.
The result is a crisis hiding in plain sight: the monitoring systems protecting some of the most valuable computing infrastructure in the world are fundamentally inadequate for the loads they are now being asked to watch.
The Numbers That Define the Problem
Traditional enterprise servers typically draw 200 to 400 watts per server, with utilization patterns that are relatively smooth. Modern AI accelerator nodes are categorically different:
-
A single NVIDIA H100 GPU can consume up to 700 watts.
-
An 8-GPU DGX H100 system draws up to 10.2 kilowatts.
-
A 42U rack of dense GPU compute can exceed 100 kilowatts.
-
A large-scale AI training cluster may pull 10 to 30 megawatts from a single facility.
Beyond raw wattage, AI workloads create electrical stress patterns that are qualitatively different from traditional compute. Training jobs create sharp, synchronized load swings as GPU clusters ramp up and down simultaneously. Inference workloads create irregular, bursty demand. The net effect on power distribution equipment is thermal cycling of a kind that decades of electrical maintenance practice did not anticipate.
Why Traditional Monitoring Falls Short
Most data centers today rely on one or more of the following monitoring approaches for electrical infrastructure:
Point-in-Time Thermal Inspections
The industry standard for decades has been annual or semi-annual thermographic inspections, in which a qualified technician walks the facility with a handheld thermal camera. These inspections capture a snapshot of equipment temperature at a single moment in time.
The fundamental problem is that electrical faults do not announce themselves during inspections. A loose connection may run cool at 40% load and reach critical temperatures only when a training cluster saturates the circuit. A thermal anomaly that develops and resolves between inspections is invisible to this approach entirely.
DCIM Power Monitoring Systems
Data Center Infrastructure Management platforms typically track power consumption at the PDU, rack, and sometimes device level. They excel at capacity planning and utilization monitoring. What they do not do is monitor the physical health of the electrical distribution equipment itself: the switchgear, transformers, busbars, cable terminations, and connections that actually deliver power to the compute.
Thermal Sensors on Compute Equipment
Server and GPU thermal management systems monitor silicon temperatures and manage cooling accordingly. These are entirely internal to the compute equipment. They provide no visibility into the electrical distribution infrastructure upstream.
"The monitoring gap is not in the compute layer. It is in the electrical distribution layer between the utility feed and the GPU. That is where AI-scale loads are creating failure conditions that traditional monitoring cannot see."
What AI-Scale Electrical Stress Actually Looks Like
Understanding why AI workloads stress electrical infrastructure differently requires a basic understanding of how electrical faults develop.
Most serious electrical failures in power distribution equipment are preceded by a period of progressive thermal degradation. A loose terminal connection develops micro-arcing. A cable termination oxidizes and increases resistance. A switchgear contact begins to pit. Each of these conditions generates heat proportional to the current flowing through it.
At traditional server loads, these developing faults may produce only modest temperature rises that progress slowly over years. At AI-scale loads, the same fault dissipates dramatically more energy. Ohm's law governs: power dissipated equals current squared times resistance. Double the current through a degraded connection and you quadruple the heat generated. The fault that might have been safely detectable for months at traditional loads can reach failure threshold in days or weeks under AI workloads.
Compounding this is the thermal cycling effect. AI training jobs start and stop, creating load swings that thermally expand and contract electrical connections. This cycling accelerates the mechanical loosening of terminals and the progression of existing faults.
The Solution: Persistent, Continuous Thermal Monitoring
The answer to this problem is not a better handheld camera or a more sophisticated DCIM system. It is a fundamentally different approach: permanently installed radiometric thermal cameras that monitor electrical distribution equipment continuously, around the clock.
Persistent Far-Field Thermography (PFFT), as developed by Power Intelligence, places calibrated thermal cameras in positions that provide ongoing visibility into switchgear, transformers, cable trays, and distribution panels. The system captures radiometric temperature data continuously, not as periodic snapshots.
The Sigma Delta Tau (SDT) algorithm then processes this continuous data stream, detecting the rate, magnitude, and pattern of temperature change rather than simply monitoring absolute temperature thresholds. This approach identifies developing anomalies far earlier than threshold-based alerts, and distinguishes genuine fault conditions from normal load-driven temperature variation.
For AI data center operators, this means:
-
Real-time visibility into the thermal state of electrical distribution equipment under actual AI workload conditions.
-
Early warning of developing faults before they reach failure threshold, measured in weeks or months of lead time rather than days.
-
The ability to correlate electrical anomalies with specific workloads, identifying which training jobs or inference patterns create the greatest electrical stress.
-
A documented, continuous monitoring record that supports maintenance planning and demonstrates due diligence to insurers and auditors.
Practical Implications for Data Center Operators
If your facility is operating or planning to operate AI workloads at scale, the electrical monitoring questions worth asking are:
-
How was your electrical infrastructure rated, and what AI-scale loads are you now drawing through it?
-
When was the last thermographic inspection of your primary distribution equipment, and what load was present during that inspection?
-
What is your current mean time to detection for a developing thermal fault in your switchgear?
-
What is the cost exposure of an unplanned outage to your AI operations?
The data center industry is in the early innings of understanding the full implications of AI-scale electrical loads. The operators who move first to close the monitoring gap will be the ones who avoid the outages that others have not yet experienced.
About Power Intelligence LLC
Power Intelligence LLC, headquartered in North Carolina, has been engineering persistent thermal monitoring solutions for mission-critical electrical infrastructure since the 1990s. Born from U.S. Department of Defense research, the company holds patented Sigma Delta Tau (SDT) and Persistent Far-Field Thermography (PFFT) technologies that provide 24/7 radiometric monitoring of substations, data centers, generation facilities, airports, and industrial sites. Power Intelligence products include Neuron, PowerIntel, PowerShot, PowerVault, ScanIR, and PoleVault.
Learn more at power-intelligence.com or request a demo today.