Skip to content
Google DeepMind's AI Cyclone Model Averaged a 230 Kilometre Five-Day Track Error Against 370 for the European Ensemble in a Nature Paper, and Its Own Authors Call Rapid Intensification an Unsolved Problem
Artificial Intelligence

Google DeepMind's AI Cyclone Model Averaged a 230 Kilometre Five-Day Track Error Against 370 for the European Ensemble in a Nature Paper, and Its Own Authors Call Rapid Intensification an Unsolved Problem

Illustration by tuput

English

A Nature paper published on 6 August 2026 reports that Google DeepMind's WeatherNext Cyclones model had a mean five-day track error of 230 km on storms of 2023 and 2024, against 370 km for the European ensemble. Its 32 authors, 22 of them Google employees, put the gain at a day or more of lead time and list what the model cannot yet do.

· · 8 min read

A Nature paper published on 6 August 2026 reports that an AI model from Google DeepMind forecast the position of tropical cyclones five days ahead with a mean error of 230 km. The European Centre for Medium-Range Weather Forecasts’ ensemble system, known as ENS, averaged 370 km on the same storms, and DeepMind’s earlier GenCast model averaged 335 km. Those scores come from storms of 2023 and 2024. The authors summarise their whole evaluation, which runs through 2025, as an average lead-time advantage of one day or more over leading operational models.

The model is called WeatherNext Cyclones. DeepMind’s better-known science system, AlphaFold, is covered in our article on its Nobel Prize, and this one produces forecasts that can be scored against what happened. Thirty-two people signed the paper. It lists 22 of them as Google employees who hold Alphabet stock, and five authors have filed a patent on the core probabilistic weather model, though not on the cyclone-specific part. The rest come from the US National Hurricane Center (NHC), Colorado State University, the UK Met Office and the University of Waterloo. Nature received the manuscript on 23 December 2025 and accepted it on 24 July 2026.

What the model learned from, and what people supplied

WeatherNext Cyclones learned from two data sets. One is about 20 terabytes of global weather analyses from the European centre, the best reconstructions of the atmosphere at past moments. The other is IBTrACS, a database of nearly 5,000 observed cyclones that the paper calls expert-curated and that records each storm’s track, peak wind and the radius of its strong winds. That second file is only about 10 megabytes, and the paper says curators at operational forecast centres compile its best-track records. The model did not discover those storms. It learned from them.

It works on a weather map with grid points about 28 km apart, using six variables at 13 heights plus six at the surface, and it steps forward six hours at a time for up to 15 days. Unlike the language models behind chatbots, it predicts numbers on a map and not words. To find cyclones, the authors mark each storm’s centre as a small bump on the map during training and add layers for its strength and size. A separate rule-based tracking program reads storm positions back off the finished forecast.

Each run produces one possible future, and the system makes 50 of them, or up to 1,000. DeepMind’s blog says a 15-day forecast takes under a minute on one TPU, a Google AI chip. What such computing costs in electricity is covered in our piece on AI energy use.

How the 230 km figure was checked

The authors froze their design choices using 2022 storms, then scored the model on 2023 and 2024 cyclones. For each test year they retrained it only on data up to the end of the previous year, so no tested storm was in its training. For 2025, a version of the model was running live, and that year’s check covers only the North Atlantic and East Pacific.

The comparisons followed NHC’s rules, which count a forecast only if every model being compared produced one for that storm at that lead time. The rivals were ENS, GenCast with its intensity bias corrected, and NOAA’s HAFS-A, a high-resolution regional model whose forecasts the study uses only for the North Atlantic and East Pacific. The paper used bootstrap tests with 1,000 resamples and says almost all improvements are significant at the 5% level.

One referee, whose report Nature published without a name, noted that such a rule can favour a cautious system that misses many storms. The authors replied that a second test, which values forecasts by their usefulness for decisions and drops no storm, covers misses.

The numbers the paper reports

The five-day track error was 230 km for WeatherNext Cyclones against 370 km for ENS and 335 km for GenCast. ENS reaches a 230 km error only after about 3.75 days, which is where the authors get just over 30 hours of extra warning at that accuracy, and about 24 hours against GenCast.

For intensity, the three-day forecast of peak wind was 3.75 knots (about 7 km/h) more accurate than HAFS, which the authors say matches roughly a decade of progress in conventional models. On probabilistic intensity scores, the model cut error by more than 50% at many lead times against ENS and GenCast.

Rapid intensification means a rise of 30 knots or more in peak wind within 24 hours. In the Atlantic and East Pacific, the model raised a score that weighs correct calls against false alarms and misses, the critical success index, from just under 0.3 to about 0.5. The authors call rapid intensification a key unsolved problem and write that no single model excels at it.

In a simulation with weights tuned on 2022 storms, adding the model to NHC’s track consensus, an average of several models, improved that consensus by 18% to 38% depending on lead time, 28% on average. The intensity consensus improved by 5% to 14%, 6% on average. The authors write that this shows conventional models still add substantial value for intensity.

What the National Hurricane Center’s 2025 scorecard says

NHC’s verification report for 2025, dated 30 March 2026, gives the live test an official reading. For the Atlantic, it says, official forecasts generally beat most models, though the Google DeepMind ensemble mean (labelled GDMI) did slightly better at short range, outperforming NHC’s track forecasts from 12 to 72 hours. In the East Pacific, GDMI beat the official track forecast and every consensus aid from 48 to 120 hours, and it was the top individual intensity model from 12 to 48 hours. The model was missing early in the season, so those comparisons leave out the first two Atlantic storms and the first five East Pacific storms.

The report also says Atlantic intensity errors were above the five-year average during the season’s many rapid intensification cases, and that official forecasts did better on them than model guidance. CNN reported on 8 October 2026 that NHC forecast Hurricane Melissa in 2025 as a Category 5 storm while it was still an 80 mph Category 1, based largely on DeepMind’s projections. NHC Director Michael Brennan told CNN the AI models have shown more consistency from one run to the next.

The limits the authors state

The paper says the model depends on high-quality analysis data and that future systems should use raw observations. It does not yet forecast rain, storm surge or wind gusts. The intensity ensemble is slightly under-dispersed, meaning its members spread less than the real errors do, and its wind probabilities are slightly overconfident at higher thresholds. At long lead times it scores worse than NHC’s intensity consensus. The supplement adds that comparisons with HAFS on intensity are more mixed, with more frequent cases where HAFS had the lower error.

A forecaster’s cautious reading

Michael Lowry, hurricane specialist at WPLG-TV in Miami and a former senior scientist at NHC, wrote on 10 September 2026 about the 2026 season so far. Using verification data from James Franklin, a former NHC branch chief, he reported that DeepMind had the lowest track error of any publicly available guidance in the Atlantic, eastern Pacific and central Pacific, up to 30% ahead of NHC at five days. Lowry is not an author of the Nature paper.

On intensity, Lowry wrote, DeepMind kept pace with NHC but did not pull ahead, and a consensus aid called HCCA did best. On Hurricane Lowell, he said, the model’s track forecasts “did initially jump around erratically” while NHC’s own shifted gradually, and he noted that NHC “knows it will sacrifice some forecast accuracy for consistency.” His numbers cover one season to date.

Matt Lanza, managing editor of the weather blog The Eyewall, told ABC News in December 2025 that the model was strong on rapid intensification but that “events outside the bounds of what’s expected will occur” and that “traditional physics-based modeling may remain essential.”

Where the Bay of Bengal fits

The Nature paper opens with the human cost of cyclones, more than 700,000 deaths and US$1.4 trillion in losses over 50 years, citing the World Meteorological Organization (WMO). WMO’s atlas counts 779,324 deaths from tropical cyclone disasters between 1970 and 2019. Bangladesh recorded 467,487, Myanmar 138,909 and India 46,784, which together make 84% of the world total.

The Nature article itself never names the North Indian Ocean, which covers the Bay of Bengal and the Arabian Sea, and nowhere mentions India’s weather department. It says results are consistent across ocean basins and points to extended data and the supplement. The supplement states that there is no clear dependence of skill on ocean basin or distance from land, and it labels one panel of an intensity-error map North Indian. NHC’s records hold HAFS forecasts only for the Atlantic and East Pacific, so other basins have no regional-model comparison.

One referee asked whether training on the Atlantic’s high-quality data risks wrongly transferring to North or South Indian Ocean storms, and noted that IBTrACS leans on US-reported values, which affects data quality outside US basins. The authors replied that the atmospheric inputs are of uniform quality worldwide, that IBTrACS also holds WMO-standard reports from every basin, and that multiple agencies had asked them about the US-values point. The 2025 live-season check was limited to the Atlantic and East Pacific because, the authors wrote, the 2025 data for the non-NHC basins had not been finalized.

The code and weights for the paper’s models are on GitHub under google-deepmind/weathernext. The repository says the models have not been endorsed by any government meteorological agency.

Share
Copied!

Sources & further reading

  1. Alet and others, Operational tropical cyclone forecasting with AI, Nature 657, 680 to 688 (published 6 August 2026)
  2. Supplementary Information for the Nature paper (extended results, basin maps, methods)
  3. Peer Review File for the Nature paper (referee reports and author replies)
  4. Google DeepMind: AI model achieves breakthrough in forecasting cyclones (6 August 2026)
  5. Google DeepMind: weathernext code and model weights (GitHub repository)
  6. Alet and others, Skillful joint probabilistic weather forecasting from marginals (arXiv:2506.10772, 12 June 2025)
  7. National Hurricane Center: Forecast Verification Report, 2025 Hurricane Season (Cangialosi, Martinez, Reinhart and Papin, 30 March 2026)
  8. World Meteorological Organization: Atlas of Mortality and Economic Losses from Weather, Climate and Water Extremes (1970 to 2019), full PDF
  9. Michael Lowry, Eye on the Tropics: The AI Hurricane Model that Continues to Blow Away the Competition (10 September 2026)
  10. Michael Lowry, Eye on the Tropics: about the author
  11. ABC News: Can AI help forecasters better predict destructive hurricanes? (6 December 2025)
  12. CNN via ABC 17 News: Why the National Hurricane Center bet big on AI to forecast Isaias (Andrew Freedman, 8 October 2026)

Researched and written with the help of AI tools and edited for accuracy. Provided for general information and discussion only, not professional advice. See our editorial standards and disclaimer. Spotted an error? Tell us.

#weathernext cyclones#google deepmind#tropical cyclones#ai weather forecasting#national hurricane center#rapid intensification#bay of bengal#nature

Enjoyed this? Get the next one.

One good read at a time, straight to your inbox. No spam, unsubscribe anytime.

More in Artificial Intelligence
The People Paid to Keep AI Safe Just Quit to Watch It From the Outside
Three safety researchers just quit two of the world's biggest AI labs to watch the industry from outside instead. Microsoft answered the AI slowdown debate with a public rulebook. Two dozen Fields medalists said AI companies are moving too fast for math to check. And India's Dhiraj Bommadevara won a gold no Indian man ever had.
Six Chinese AI Firms Got Caught Copying America's Best Models. The US Fix: Quietly Make Their Answers Worse.
A US security advisory accuses six Chinese AI companies of mass copying American models and recommends quietly serving them worse answers. Plus DeepMind's atlas of every possible human DNA mutation, Apple's AI-first iPhone launch, and India's biggest rail expansion in years.
Scale's SWE-Bench Pro Board Puts the Best Coding Agent at 61.5%, Maintainers Reject Many Fixes That Pass the Tests, and METR Says Its Latest Trial Cannot Show How Much Time Agents Save
Benchmarks say agents are strong. Maintainers, a randomised trial and a pile of incident reports say the picture is narrower. Here is what each number actually measures.
← all articles