Google DeepMind released WeatherNext 3 on September 3, 2026, and the press release leans hard on marketing superlatives. But buried in the accompanying arXiv paper (Rasp, Babenko, Masters, Battaglia, and 30+ authors) is a model redesign that is genuinely worth examining for anyone who trains or deploys diffusion-free probabilistic weather models. The primary announcement itself is short on technical detail; the value here is in the paper's architecture, benchmark, and artifacts sections. This is not a rewrite of the announcement. This is a breakdown of the numbers, and the parts of the paper that the blog post quietly leaves out.
I read the Google announcement, the arXiv paper itself, and Google's developers research page for benchmark methodology. The gap between them is where the interesting engineering lives.
What WeatherNext 3 Actually Is
WeatherNext 3 is the third generation of DeepMind's AI weather forecasting stack, sitting atop GraphCast (deterministic), GenCast (probabilistic diffusion), and WeatherNext 2 (the FGN model this one replaces). The short version: WeatherNext 2 was trained only on NWP-generated analysis data. WeatherNext 3 is the first in the lineage trained to ingest raw satellite observations directly and to predict satellite-derived precipitation, tropical cyclones, and sparse station readings as first-class outputs.
That sounds incremental. It is not. The paper makes a sharper claim: WeatherNext 3 stops treating weather forecasting as three separate stages (data assimilation, forecasting, post-processing) and runs them as one continuous function. That single architectural decision is worth unpacking because it changes what the model can learn.
Architecture In Detail
WeatherNext 3 is built on the same encode-process-decode structure as WeatherNext 2, using what DeepMind calls Functional Generative Networks, or FGN for short. The model lives on an icosahedral processor mesh, which sidesteps the pole singularity that plagues latitude-longitude grids. Between WeatherNext 2 and WeatherNext 3, two capacity numbers went up: the latent width grew from 768 to 1024, and the mesh transformer deepened from 24 layers to 32.
Here is the structural change that matters most. WeatherNext 2 was a single-modality model, trained only on analysis variables at 0.25 degree resolution. It generated output every 6 hours, then used a separate lambda upsampler model to fill in the hourly frames you actually see in the user interface. WeatherNext 3 predicts hourly resolution natively. It learned to interpolate time internally instead of bolting on an upsampler as a second model.
The model handles three resolutions at once, each mapped to the shared mesh by its own encoder and decoder so no data has to be regridded to a single common grid.
| Dimension | WeatherNext 2 | WeatherNext 3 |
|---|---|---|
| Latent width | 768 | 1024 |
| Transformer depth | 24 layers | 32 layers |
| Surface temp resolution | 0.25 deg (~25 km) | 0.05 deg (~5 km) |
| Surface variable resolution | 0.25 deg | 0.1 deg (~10 km) |
| Atmospheric resolution | 0.25 deg | 0.25 deg (~25 km) |
| Temporal cadence | 6-hourly + upsampler | hourly (native) |
| Ensemble members | 64 | 64 |
| Satellite input | indirect | 11-channel geostationary mosaic, direct |
| Precipitation targets | analysis-based | satellite-derived (PARDIG, IMERG) |
| Station output | no | continuous query head |
The inputs table in the paper is where the redesign becomes concrete. WeatherNext 3 ingests ERA5 back to 1959, HRES analysis, an 11-channel geostationary satellite mosaic at hourly cadence, its own satellite-derived precipitation reanalysis called PARDIG, NASA's IMERG precipitation product, and sparse in-situ station readings for 2-meter temperature and dewpoint. It predicts all of that plus tropical cyclone tracks as a discrete off-grid tabular output.
The most interesting piece is the station head. Instead of predicting on a fixed grid, WeatherNext 3 decodes predictions as a continuous query. It interpolates the latent to any latitude and longitude you ask for, combines it with elevation, land-sea mask, and time-offset metadata, and outputs temperature there. That means you can ask for the temperature in a specific valley or off any coast, at any hour, without the model being pinned to a predefined grid.
Benchmark Performance
The numbers the paper reports are strong, but the evaluation design deserves scrutiny. WeatherNext 3 is measured against WeatherNext 2, against ECMWF's operational systems (HRES and ENS), and against the competing AI model AIFS ENS v2. Let me separate what is meaningful from what is noise.
On upper-level variables, WeatherNext 3 improves CRPS by roughly 5% over WeatherNext 2. That sounds small until you translate it into lead time. The paper notes a 5% improvement roughly corresponds to about 6 hours of additional forecast skill at the same accuracy level. To put that pace in context, operational weather forecast skill has historically advanced by roughly one day of lead time per decade since the numerical weather prediction revolution documented by Bauer, Thorpe, and Brunet in 2015. DeepMind bought six hours of new skill in about a year of model iterations. That is not incremental. It is an order of magnitude faster than the historical trajectory.
The precipitation results are the headline. Against IMERG, WeatherNext 3 cuts CRPS by up to 60% over the previous best models, 30% against MRMS radar reanalysis, and 10% against rain gauges for early lead times. Precipitation is the classic failure mode for AI weather models because rain is driven by convective processes at scales that coarse grids cannot resolve. Training directly on satellite-derived precipitation rather than on analysis fields is what unlocks this.
The station head is equally strong. Against weather stations the model never saw during training, it reduces 2-meter temperature CRPS by up to 30% over WeatherNext 2 and up to 40% against ECMWF ENS. That last number is the one that should make traditional forecasters uncomfortable: the improvement holds against a held-out global station network, not just against reanalysis.
The real-time evaluation is the most honest part of the paper. DeepMind ran a quasi-real-time comparison over the six weeks between July 1 and August 11, 2026, using the actual operational data feeds. WeatherNext 3 beat AIFS ENS by about 10% in the first forecast week on upper-level variables. The authors correctly flag that six weeks is a small sample and that WeatherNext 3 and AIFS ENS both include some of the same training data, which they call 50r1. They do not pretend that confound is gone. That level of disclosure is rare in model release papers.
Engineering Perspective: The Design Trade-Offs
Every choice here is a trade-off, and the paper is unusually candid about the costs.
The continuous station head introduces a real artifact problem. Because the model optimizes marginal CRPS at each query point independently, it can cheat on the covariance structure. The paper shows visible hexagonal patterns in precipitation and station output, a direct reflection of the underlying icosahedral mesh. These patterns are most severe in individual ensemble members and fade in the ensemble median. There are also temporal discontinuities across the 6-hour outer time-step boundaries, visible as jumps in the hourly station output. The authors note that the station head exhibits a per-ensemble-member global bias that shifts every 6 hours as a new noise vector is drawn.
This is the honest trade-off: you gain arbitrary-location, arbitrary-time query flexibility, and you pay for it with covariance artifacts that only smooth away in aggregation. For downstream applications that rely on median or threshold-exceedance probabilities, the forecasts remain usable. For applications that need pixel-level joint structure, the artifacts matter.
The model also got bigger, and bigger means more compute. WeatherNext 2 generates its full 64-member ensemble in a single forward pass, which is roughly 8x faster than the iterative diffusion approach GenCast used before it. WeatherNext 3 adds capacity on top of that single-pass ensemble, so it is still dramatically cheaper to run than a physics-based NWP system, but the 1024-latitude, 32-layer model is not free. DeepMind had to introduce spatial sharding of the processor mesh to fit it on available accelerators. The paper does not publish the exact token count or training FLOPs, which is a genuine gap for anyone trying to budget inference costs. For a broader look at how the data center GPU landscape is reshaping around single-purpose AI inference, the 2026 GPU comparison covers the same hardware these models end up running on.
Uncertainty modeling is handled through functional perturbation ensembles that capture both aleatoric and epistemic uncertainty. WeatherNext 2 used four independent model seeds for the epistemic component. WeatherNext 3 drops to two seeds and instead adds what the paper calls "epistemic dropout," which frames dropout as sampling from an exponential family of networks. This is a clever answer to a real problem: a bigger model with fewer seeds tends to overfit and under-spread its ensembles, and dropout variance is the counterweight.
Why This Matters
For builders, three things stand out.
First, the end-to-end consolidation is the real innovation. Traditional weather forecasting splits assimilation, forecasting, and post-processing into separate pipelines, each with its own assumptions and error budgets. WeatherNext 3's station head eliminates the need for a separate station-calibration stage entirely, and the native hourly resolution removes the upsampler. That is not just cleaner engineering. It removes a stage where biases get introduced, and it shaves latency. For an operational system, removing a post-processing step that used to run on top of the forecast is worth more than the raw accuracy delta.
Second, the satellite input changes the economics of forecasting in underserved regions. WeatherNext 3 was specifically designed to help Latin America, Africa, and the Asia-Pacific, regions that historically get poor high-resolution forecasts because running regional NWP there costs enormous supercomputing budgets. A model that ingests satellite mosaics directly and queries any point on the globe does not need a regional supercomputer. That is a genuine democratization of forecast quality, and it is why the press release emphasizes those regions so heavily.
Third, the model is exposed as queryable data, not just static grids. WeatherNext 3 output is available in BigQuery and Earth Engine and bulk-downloadable from Google Cloud Storage in Zarr format. The continuous station head means a developer can pull 2-meter temperature for a specific wind farm site at turbine height rather than the nearest grid point. For the clean energy use case the paper calls out, predicting 100-meter wind speeds, cloud cover, and solar radiation, that granularity directly affects how accurately a grid operator can match renewable generation to demand.
The Honest Gap
No release is clean, and this one is refreshingly candid about its flaws. Beyond the hexagonal mesh artifacts and the 6-hour temporal jumps, the paper does not publish inference compute costs, exact parameter counts, or the total training budget. The real-time evaluation window of six weeks is honest but thin. And the paper admits that at 6-hour and occasionally 12-hour lead times, a subset of variables actually regress slightly against WeatherNext 2, with no definitive explanation yet. These are not deal-breakers, but they are the parts a specialist would probe first.
Our Read
WeatherNext 3 is less about a single record on a leaderboard and more about changing the shape of the forecasting pipeline. The 5% upper-level improvement translating to six hours of new skill, the 60% precipitation CRPS gain against IMERG, and the removal of two post-processing stages together represent a genuine architectural step, not a tuning bump. The artifacts are real but manageable for aggregate statistics. Whoever needs pixel-level joint precision should wait for the next iteration. For hourly operational forecasting, clean energy planning, and high-resolution output in underserved regions, WeatherNext 3 is already one of the strongest systems available, and it runs in a single forward pass.
For anyone watching how AI models move from research papers into production infrastructure, this is a useful case study in consolidating a multi-stage pipeline into one differentiable function and paying for it with artifacts instead of latency.
References
- Introducing WeatherNext 3, our most advanced and accurate global weather AI model. Google. The primary announcement covering the headline resolution and frequency figures.
- WeatherNext 3: Increasing resolution and performance of global weather models with raw observations. arXiv. The full paper with architecture, training curriculum, and benchmark tables.
- Research and benchmarks | WeatherNext. Google for Developers. Independent methodology page covering WeatherBench 2 and Brightband evaluations.
- The quiet revolution of numerical weather prediction. Bauer, Thorpe, Brunet, Nature 2015. Cited in the arXiv paper for the historical one-day-per-decade skill trajectory.
- Operational tropical cyclone forecasting with AI. Alet et al., Nature 2026. Cited for the cyclone evaluation protocol.