Frequent Off-Grid Power Failures at Hyperscale Data Centers, with Downtime Costs Reaching Tens of Millions of Dollars
Nashnova编辑部
Three of the four off-grid data centers operating in the U.S. have already suffered equipment failures. Violent power swings from AI workloads are destroying generation hardware — a single day of downtime may cost tens of millions of dollars, putting the entire off-grid model's viability in question.
Four off-grid data centers — what went wrong at three of them?
The U.S. currently has four off-grid or partially off-grid data centers in operation: xAI's Colossus and Colossus II, Crusoe's Stargate Abilene, and Vantage's VA 2. Three have experienced documented failures.
Failures include gas engine crankshaft fractures, gas turbine cracks, and a critical power fault that forced a switch to diesel backup for roughly a day.
This means → off-grid power is not a case of one unlucky project. A three-out-of-four failure rate points to a systemic technical flaw.
Why does AI load wreck power-generation equipment?
AI workloads produce wildly uneven power curves: demand spikes during computation and drops sharply when results converge. In plain terms = generators are pushed to full output one moment, then left nearly idle the next — repeated "emergency braking" oscillations accelerate mechanical wear.
As data centers scale, these oscillations can reach hundreds of megawatts or even gigawatt-level swings. Energy consultancy Wood Mackenzie warns that such violent fluctuations can directly cause engine and turbine shaft fractures.
Off-grid systems face an additional challenge: equipment heterogeneity — power plants are cobbled together from different vendors' hardware, making coordination difficult. One hyperscaler demanded power response times of tens of milliseconds, versus the industry norm of hundreds. This reflects a reality where AI's demands on power systems already far exceed the design envelope of existing equipment.
One day of downtime — who loses how much?
Anthropic pays xAI $1.25 billion per month for compute on Colossus and Colossus II. At that rate, a single day of downtime could cost tens of millions of dollars.
For xAI, it is a direct revenue loss. For Anthropic, it is the opportunity cost of lost compute — interrupted model training, delayed product delivery.
This means → the cost of an off-grid failure is not just the repair bill. Both sides of the deal bleed simultaneously: the power provider pays penalties, the data center loses revenue, and the compute buyer's business stalls.
Can the power providers absorb the hit?
Industry practice: providers do not charge when they cannot deliver power. Some contracts include penalty clauses for non-delivery, and repair and replacement costs typically fall on the provider.
Electrical engineer Nina Sadighi disclosed that one of her clients' off-grid power systems suffered multi-day downtime, forcing the provider to replace equipment worth several million dollars.
Providers like Solaris have begun inserting liability caps into contracts. In plain terms = they know this business can generate outsized losses and are trying to limit exposure through contract terms — but if equipment degradation continues to outpace projections, margins will keep eroding.
What does this mean for the off-grid model?
Big tech chose off-grid power to bypass grid permitting bottlenecks and accelerate AI infrastructure deployment. But the current failure rate shows that technical maturity has not kept pace with the speed of expansion.
This reflects a deeper tension: the growth rate of AI compute demand is forcing energy infrastructure into a "deploy first, fix later" mode — precisely the operating philosophy that power systems are least forgiving of.
For investors, the core risk in off-grid power is no longer "can it be built?" It is "once built, can it run reliably?"
Content is for reference only, not financial advice.