SDSignal Desk

Update to Google’s AI weather model improves forecast accuracy

Sep 8, 2026, 11:00 AM · Ars Technica

Image: Ars Technica

WeatherNext 3 adds raw satellite inputs and hourly runs, cutting reanalysis lag—and posting measurable gains over WeatherNext 2 and ECMWF’s AI model, with a few odd artifacts.

Why it matters

Ars Technica reports Google released WeatherNext 3, detailed in a white paper, with the biggest change being ingestion of some satellite weather data. That shortens the lag between current conditions and a new forecast. AI weather models already match much of traditional-model skill at far lower compute cost, which means they can be run more often.

Most AI weather systems have trained and inferred almost entirely on global reanalyses—consistent atmospheric snapshots that fill gaps where measurements are missing, typically every six hours. Raw observational detail can be lost in that blender. WeatherNext 3 moves closer to traditional practice by adding satellite observations and raising forecast frequency to hourly, while also increasing spatial resolution and model size (with process tweaks to contain compute).

Google says the model is now the forecast source across Search, Gemini, and Maps. Performance claims versus WeatherNext 2 include roughly 5% better upper-atmosphere accuracy—about six more hours of accurate lead time—and up to 30% better surface-temperature accuracy via a location-aware land/ocean/elevation calculation trained on station data. Results generally beat the ECMWF AI model on the paper’s metrics, with a curious early-horizon exception.

The Signal Desk read

Signal Desk’s read: the story that matters is not “AI weather got a bit better”—it is Google quietly closing the architectural gap that kept ML forecasts dependent on reanalysis cadence while traditional models ate raw observations.

Hourly satellite-informed updates are the operational unlock. If AI models can refresh as often as compute allows without a six-hour reanalysis bottleneck, they stop being clever offline demos and become the default layer inside consumer products—which is exactly where WeatherNext 3 already sits. A separate ML model trained on satellite precipitation estimates, yielding multiple precip forecasts, plus pointwise surface temperature and dew point that know land versus ocean and elevation, are incremental physics-adjacent crutches on an otherwise black-box pattern machine. That hybrid instinct is healthy: pure pattern matching hits ceilings that sparse physical conditioning can raise.

What the white paper under-plays—and Ars rightly flags—are the oddities. For several variables, WeatherNext 3 looks worse than peers on the initial six-hour-ahead comparison before pulling ahead across a 15-day horizon. Precipitation maps can show hexagonal grid-shaped blobs. Ensemble-style surface-temperature snapshots sometimes shift global average temperature up or down instead of keeping the global mean stable while varying locally. Those are the kinds of artifacts that matter more for professional uptake than a headline 5% upper-air gain.

The likelier read is that Google is optimizing for productized, frequently refreshed forecasts at Google scale, not for replacing every ECMWF operational workflow tomorrow. Expect competitors to copy the raw-observation diet. Treat the “major step forward” language as directionally fair on latency and data richness, and treat the early-horizon and grid-artifact issues as open research debt, not footnotes.

Context

AI weather models from Google and peers have competed on skill-per-compute against classical physics simulations. Reanalysis dependence was a shared limitation; WeatherNext 3’s satellite path and hourly cadence are the explicit answer. ECMWF’s own AI model is the paper’s main external yardstick.

Who feels it

Consumers using Google Search, Maps, and Gemini
Forecast cards and assistant answers now ride WeatherNext 3; accuracy and update frequency changes will show up in everyday weather UX before most users notice the model name.
Meteorology and energy desks
Watch early-horizon skill and precipitation morphology before swapping traditional guidance; hexagonal artifacts and global-mean drift in ensembles remain caution flags.
Competing AI weather labs
Pressure rises to pull in low-latency satellite streams rather than relying on reanalysis alone.
Google Research / DeepMind weather teams
Need transparent explanation of the six-hour underperformance dip and grid/ensemble oddities if they want operational credibility beyond Google’s own surfaces.

What to watch

  1. Whether follow-on papers or ops notes explain the worse-then-better six-hour skill pattern.
  2. Independent bake-offs of WeatherNext 3 versus ECMWF AI and classical models on severe-weather cases.
  3. If hexagonal precip artifacts and ensemble global-mean drift shrink in the next revision.

Read the original

Continue at the source.

Ars Technica

Companies: Google