Standard testing methods create a false impression that artificial intelligence can reliably forecast electric power outages. While high test scores suggest these systems are ready to protect infrastructure, they collapse when facing new weather events or unmapped regions. The failure happens because standard evaluations split data randomly, allowing models to cheat by peeking at nearby geographic points and adjacent timestamps.
Weather conditions and power grid lines share strong physical links across space and time. A model tested on random data splits operates like a student copying answers from a neighboring desk rather than solving the math. When a storm strikes an unseen area, this shared information vanishes and the algorithm cannot connect weather forces to electrical damage. Adding embeddings from the Prithvi WxC foundation model produced only minor and inconsistent spatial gains without solving failures during new storm events.
Researchers tested outage prediction models using public United States East Coast records between 2018 and 2023. They evaluated the systems under strict conditions that held out entire states or entire storm events from the training data. Under these realistic deployment tests, predictive accuracy dropped so sharply that the systems frequently failed to beat a simple null baseline.
Developing operational outage tools requires expanding real-world data coverage and enforcing strict state-level and event-level evaluation protocols. Researchers indicate that progress depends on fixing structural data constraints instead of pursuing marginal tweaks to machine learning algorithms.
