Executive Summary

For the first time, eight AI climate models from six research groups stood side by side on a single physics baseline. This is AIMIP Phase 1, led by the Allen Institute for AI. The result comes down to one sentence: they got the past average right, but half of them missed the trend heading into the future. This report does not ask whether AI is good at climate. It takes up the prior question — that before you can answer that, you first have to line up scattered models against the same yardstick.

In reproducing the training period, the best AI models cut the time-mean error of surface temperature to roughly half that of the physics model — and did so almost regardless of architecture. Yet when they were tested on a period they had never seen, which happened to be the hottest decade in the observational record, they split: some tracked the warming trend accurately while others significantly underestimated it. It is a fracture the average, taken alone, would never have shown.

For Pebblous readers, this is data quality one level up. Beyond the freshness of the input data, it is a story about the quality of the evaluation data and the evaluation protocol. Without a standard, there is no comparison; without comparison, there is no trust.

Mean-climate error

How much the best AI models cut time-mean error against the physics baseline (GFDL-CM4)

8

Models on one yardstick

AI climate models from six groups, compared side by side for the first time

2015–2024

Held-out test period

A decade the models never trained on — and the hottest decade on record

~0.21°C

Real warming per decade

ERA5 observed rate (1979–2025) — the "true number" half the models undershot

1

Eight Models, One Yardstick

Over the past few years, AI weather and climate models have poured out at a remarkable pace. Google DeepMind's GraphCast, NVIDIA's FourCastNet, Google Research's NeuralGCM, ECMWF's AIFS — each arrived with a report card claiming it had "beaten the physics model." But those report cards were scored on different data, over different evaluation windows, against different metrics. We had a set of champions who had each won their own league, and no one had ever put them on a single track to run side by side.

In May 2026, the Allen Institute for AI (Ai2) laid down that track for the first time. AIMIP (the AI Weather and Climate Model Intercomparison Project) Phase 1 placed eight AI simulations from six research groups onto a single baseline — NOAA's physics-based coupled model GFDL-CM4 — and compared them against the same data, the same protocol, the same yardstick. It is the first multi-party intercomparison to line up AI climate models against one another.

The word "intercomparison" matters here. This is not the work of digging into why any single model is wrong; it is the work of asking whether these models can be fairly placed side by side in the first place. That is where it parts ways cleanly with earlier benchmarks. The table below lays out how AIMIP differs from the infrastructure that came before it.

Infrastructure What it compares Character
AIMIP Multiple AI climate models against each other, on one physics baseline Multi-party intercomparison — this study
WeatherBench2 Mainly short-range weather forecast accuracy Established forecast benchmark (3+ years running)
ClimateBench Emulation of a single physics model (NorESM2) Single-model imitation — not an intercomparison

The very youth of this benchmark is the hook of the story. AIMIP's code repository is new infrastructure, released alongside the paper, with GitHub stars in the dozens. WeatherBench2, running for more than three years, is in the six hundreds; GraphCast, a single-model repository, reaches into the thousands. That gap is not a mark of inferiority — it is because the problem of "lining up multiple AI climate models" started far later than short-range forecasting, and is only beginning now. In other words, until this moment we had no common yardstick to compare AI climate models fairly at all.

2

Designing the Yardstick

A benchmark earns trust from its design, not from a flashy headline result. The yardstick AIMIP built rests on three questions: what it compared (the participating models), how it split the data (training vs. validation periods), and what it measured (five evaluation axes).

2.1 What it compared — eight models and one baseline

The entrants were eight simulations from six research groups, differing in architecture and resolution alike. Alongside them stood NOAA GFDL's physics-based coupled model, GFDL-CM4, as the baseline — because to claim AI has "beaten" physics, you have to run it against that physics model under the same conditions.

Model Submitting group Grid resolution
ACE2.1-ERA5 Allen Institute for AI 1°×1°
ArchesWeather Google DeepMind / INRIA 1°×1°
ArchesWeatherGen Google DeepMind / INRIA 1°×1°
cBottle-1.3 NVIDIA HEALPix ~0.9°
DLESyM U. Washington / NVIDIA HEALPix ~0.9°
MD-1.5 v0.9 U. Maryland Submitted at 1°×1°
NeuralGCM Google Research Submitted on 7 pressure levels
NeuralGCM-HRD Google Research 1°×1°
GFDL-CM4 (physics baseline) NOAA GFDL Physics-based coupled model

The paper does not disclose parameter counts per model. "Six research groups" counts the groups that submitted models; the paper's full list of co-authoring institutions is larger.

2.2 How it split the data — time it learned, and time it didn't

The heart of the design is how it cut time. Every model trained on 1979–2014 reanalysis data — the satellite era — and was then validated on 2015–2024. That validation window is a held-out period the models never saw during training. It is a real future, not a simulated one. The data follows the NetCDF and CF standards and is released under CC-BY 4.0, so anyone can re-grade the models under the same conditions. Reproducibility is what qualifies something as a benchmark.

One asymmetry is worth a footnote. Unlike the others, DLESyM trained on 1983–2016 and on only a subset of variables. For a fair comparison, differences like this should themselves be recorded and surfaced. That is precisely this report's meta-point.

2.3 What it measured — five evaluation axes

Rather than collapse performance into a single score, AIMIP broke it into five axes — to catch the tail that disappears when you look only at the average. These five axes are both a map of the story to come and a checklist a company can use, as is, when vetting any AI model.

E1 · Bias

How accurately it reproduces the mean climate of the training period.

E2 · Trend

Whether it tracks the long-term warming trend or underestimates it.

E3 · ENSO response

Whether it responds correctly to large-scale variability like El Niño and La Niña.

E4 · Temporal variability

Whether it preserves the true range of day-to-day temperature swings.

E5 · Out-of-distribution generalization

Whether it responds in line with physical law even outside the training distribution (+2°C / +4°C forcing).

3

They Nailed the Average

Once the yardstick was in place, the first result came through sharply. On the bias axis (E1) — the ability to reproduce the mean climate of the training period — the best AI models beat the physics model. And not by a hair.

The best AI models cut the time-mean error of fields like surface temperature to about half that of the physics baseline, GFDL-CM4. And this edge was not a fluke of one architecture. Diffusion model, graph neural network, or hybrid with differentiable physics: almost regardless of architecture, the AI side drew the mean climate better.

Physics baseline (GFDL-CM4) 1.0× Best AI models ~0.5× Error 0 Baseline error = 1.0

Conceptual schematic — relative value with the physics baseline's (GFDL-CM4) time-mean error set to 1.0. The paper presents exact per-field absolute errors as figures, not a table.

It was not just the mean climate. On ENSO response (E3), which deals with El Niño and La Niña, most AI models actually produced slightly smaller errors than GFDL-CM4. Even on the hard problem of large-scale variability, AI edged past the physics model on average. Stop here, and the story looks set to end with "AI has conquered climate."

But the average is a trap. Getting the average right means drawing "how things usually are" well — not "which way things are changing," nor "the rare but decisive value." In fact, the same data already showed the seeds of a fracture. Precipitation bias varied widely in sign and location across models, especially near the equator, where even physics models struggle. And on day-to-day temperature variability (E4), nearly every AI model tended to underestimate the range compared with observations (ERA5). This was the one axis of the five for which the paper left an actual percentage — on the order of 2–10% of the ERA5 variability magnitude. A habit of painting the world a little calmer than it really is — a habit that grows into a much bigger problem in the next chapter.

4

They Missed the Trend

Moving into the validation window (2015–2024), the picture changed. On the trend axis (E2), the eight models split in two. Some tracked the last decade's warming accurately; others significantly underestimated it. Models that had performed shoulder to shoulder over the training period revealed different faces for the first time when facing a future they had not learned. The schematic below shows that split conceptually.

Temp. Training 1979–2014 Held-out 2015–2024 Observed (real warming) Trend-tracking model Underestimating model Widening gap Nearly overlapping

Conceptual schematic — the paper presents exact °C/decade values in a time-series graph (Figure 3), not a table. This shows direction only; the tick values are not illustrative figures.

Two things need to be stated clearly here. First, this is not "a failure of AI in general." The trend-pattern error of the models that tracked the trend well (NeuralGCM, the ArchesWeather family, DLESyM) was on par with GFDL-CM4 — meaning the best AI matched the physics model on trend, too. The problem is confined to some of the eight, about half. Second, the failure did not appear out of nowhere in the held-out window. Most of these models were already, from the training period (1979–2014), underestimating the trend a little; the gap simply widened in the validation window. The "habit of painting the world a little calmer" from the previous chapter accumulated over time.

Models that tracked the trend

NeuralGCM · ArchesWeather · ArchesWeatherGen · DLESyM — captured the observed warming relatively accurately

Models that underestimated it

ACE2.1-ERA5 · MD-1.5 v0.9 · cBottle-1.3 — significantly undershot the last decade's warming

4.1 Throwing it a physically impossible scenario

To check whether a model that missed the trend truly understood physics, you have to throw it an extreme that isn't in the training data. AIMIP ran a physically impossible shock test (E5): instantaneously raising sea surface temperature by +2°C and +4°C. Because no such stimulus exists in nature, this is where you find out whether a model responds from the statistics it learned or the physics it learned.

Among the models the paper describes, responses split two ways. NeuralGCM-HRD and DLESyM showed the physically expected pattern of generalized warming. ACE2.1-ERA5, cBottle-1.3, and MD-1.5 v0.9, by contrast, responded to a warming-ocean forcing by cooling the land — a response whose very direction contradicts physical law. Not merely wrong in magnitude, but wrong in sign. That the three models that underestimated the trend and the three that broke physics here line up so cleanly is no coincidence.

The directional verdict on the shock test is limited to the five models the paper explicitly describes. It is not generalized to a verdict on all eight.

4.2 The hottest decade, of all decades

That the validation window is 2015–2024 is the climax of this story. This decade is the hottest on record. The ten warmest years ever measured all fall within it. The global warming rate, per ERA5, is about 0.21°C per decade over 1979–2025 — and the more you narrow the observation window toward the present, the steeper that value gets. In other words, 2015–2024 was an inflection stretch where the trend itself curved upward.

The very window where AI models failed to capture the trend coincides exactly with the window where the planet warmed most dramatically. That coincidence was a stroke of luck for the benchmark. The moment a standard was set, that standard pointed straight at the worst-case condition. It is a fracture that, looking at the average alone, would never have surfaced.

5

The Lure of Speed, the Trust Gap

So why did these models pour out so fast? The answer is speed and cost. On this front, AI weather and climate models are attractive in a way physics models can't touch. The numbers below show the size of the temptation.

Model Speed / cost advantage
GraphCast Generates a 10-day forecast in under a minute on a single TPU (traditional NWP takes hours on a supercomputer); roughly 1,000× less energy
ECMWF AIFS Operational since Feb 2025. Up to 20% improvement over the physics model on some forecasts; roughly 1,000× less energy per forecast
FourCastNet Up to 45,000× faster than traditional NWP (the most aggressive claim)

A minute versus hours; 1,000× the energy saved. In front of numbers like these, it is hard to find a reason not to switch to AI. And in short-range weather forecasting, doing so is in fact the right call. For day-scale forecasts like a typhoon's track, AI models already match or beat physics models.

10-day forecast generation time (conceptual) Traditional physics model (NWP) Hours Supercomputer AI model (GraphCast, etc.) Under 1 min · 1 TPU ~1,000× less energy

Conceptual schematic — not to a log scale, direction only. Based on published GraphCast and ECMWF AIFS figures; actual scale differs.

The problem is that "fast and cheap" does not automatically guarantee "trustworthy." This gap is what AIMIP exposed. In the body — the average and short-range forecasts — AI is strong; in the tail — the long-term trend, extremes, and unfamiliar forcing — half of them waver. The developers know this limit themselves. The NeuralGCM paper boasts of saving computation over traditional GCMs by orders of magnitude for decade-scale climate projection, yet in the same document states plainly that it should not be extrapolated to abrupt future-climate scenarios. The model itself concedes a limit of the same grain as AIMIP's held-out failure.

Speed is a reason to adopt, but it is no basis for trust. Which AI model to adopt should be judged not by how fast it is, but by how it behaves in the very region we care about — the tail, not the body. And to look into that tail, you first have to have a yardstick that measures the tail at all.

6

From Mechanism to Method

We have written before about why AI climate models go wrong. In a piece on the cold bias of AI climate models, we dug into the mechanism by which an individual model draws certain regions and seasons colder than they really are. That was the question "where does this model go wrong?"

The question AIMIP poses sits one level up. Not "where does this model go wrong," but "can we fairly compare the wrongness of many models?" If the earlier piece was about mechanism, this one is about method. However precisely you diagnose the error of an individual model, if you can't place those diagnoses on the same yardstick, you cannot answer the practical question of "which model is better." AIMIP's contribution is not a new model — it is that scattered models were, for the first time, placed on the same scale.

And the moment they were on the scale, the tail hiding beneath the surface of the average came into view. The half that missed the trend, the shock responses that broke physics, the habit of flattening day-to-day variability. None of this is visible from individual report cards; all of it appears only when the models are lined up under a common protocol. Even the youth of the benchmark (a fledgling repository with dozens of stars) is part of the story. We are only just beginning to learn how to compare AI climate models fairly.

Without a standard, there is no comparison; without comparison, there is no trust. The proposition AIMIP made in the climate domain applies, in fact, to every domain that tries to judge AI models by performance.

7

Why This Matters to Pebblous

Pebblous watches this research not for climate itself. It is because the problem AIMIP ran into in the climate domain has exactly the same structure as the one we meet every day in the domain of industrial data and models.

Data quality, one level up

When we have talked about data quality so far, we have mostly meant the freshness, distribution, and domain fit of the input data. AIMIP points to the layer above it: the quality of the evaluation data and the evaluation protocol. How you split the training and validation periods, how you design the held-out set, what forcing you use to probe outside the distribution. When this design is weak, a flashy "half the mean error" metric masks a collapse in the trend. The distribution of the training data determines the model's internal representation (the trend was already being underestimated from the training period), and that limit is amplified in the held-out window — and that causal chain is what needs diagnosing.

How to read a vendor's report card

When a company adopts an AI model — whether for climate, demand forecasting, or process simulation — an actionable lesson follows: don't take the vendor's "X% improvement in average accuracy" at face value. AIMIP's five evaluation axes become a vetting checklist as they are. Three questions you must always ask:

  • 1Is the validation window a genuine future (held-out), or training data lightly altered and reused?
  • 2Did it pass an out-of-distribution stress test? A model strong only on ordinary data flips sign at the extremes.
  • 3How does it perform in the tail (trend, extreme values), not the average? What we care about is usually the tail, not the body.

Make things comparable first

What AIMIP did in the climate domain — making scattered models diagnosable against a common yardstick — is the same in spirit as what Pebblous sets out to do in the domain of industrial data and models: make things comparable before discussing performance. AIMIP's observation, that a model trained on the physical world responds in ways that contradict physical law outside its training distribution, shows in the language of Physical AI why designing for real-world generalization matters. Setting a standard is unglamorous, but without it every performance claim that follows stays unverifiable.

AIMIP is infrastructure that has only just begun. It has lined up eight models once, and what the next Phase will reveal is still unknown. But the difference between setting a yardstick and not setting one is already clear. Once the yardstick was set, the story beneath the surface of the average became visible for the first time. That alone is enough for this benchmark to have earned its keep.

R

References

Academic papers

Benchmarks, models, industry

  • 2.Allen Institute for AI (2026). Introducing AIMIP: The AI Weather and Climate Model Intercomparison Project. (Source for the factor-of-2 advantage in mean-climate error.)
  • 3.ai2cm/AIMIP GitHub repository (2026). Protocol, SUBMISSIONS.md, NetCDF/CF spec, CC-BY 4.0. (Training 1979–2014 / validation 2015–2024 split.)
  • 4.ECMWF (2025–2026). ECMWF's AI Forecasts Become Operational (Feb 2025) and the AIFS v2 announcement (May 2026). (Up to 20% improvement over the physics model; ~1,000× less energy.)
  • 5.google-research/weatherbench2 GitHub repository. (Contrast with the established short-range forecast benchmark.)
  • 6.google-deepmind/graphcast GitHub repository. (Contrast in single-model ecosystem scale; 10-day forecast in under a minute.)
  • 7.duncanwp/ClimateBench GitHub repository. (Single-physics-model emulation — a layer distinct from AIMIP's intercomparison.)

Observations and climate indicators

  • 8.Copernicus Climate Change Service. Climate Indicators — Temperature and Global Climate Highlights 2024. (ERA5 warming rate ~0.21°C/decade, 1979–2025.)
  • 9.NOAA National Centers for Environmental Information (2025). 2024 Annual Climate Report. (Warming rate 0.20°C/decade; 2024 at +1.35°C; ten warmest years on record = 2015–2024.)
  • 10.NASA GISS (2025). 2024 Temperature Analysis. (+1.47°C above pre-industrial.)