Ensembles
Running the model many times from slightly different starting points converts a single answer into a spread, which is the honest output.
Running the model more than once is not a bug — it is the method
A deterministic forecast gives you one answer. Run the model from today's observations, step it forward seventy-two hours, and read off the temperature. The number is precise. The precision is misleading.

The problem is that the starting point is never exactly known. Every observation carries a small error — a radiosonde drifts sideways in a crosswind, a buoy bobs on a wave, a satellite retrieves radiance and leaves the temperature to be inferred. Data assimilation stitches these imperfect measurements into the best available picture of the atmosphere, but best available is not perfect. Whatever uncertainty lives in the initial state grows as the model runs forward, because the atmosphere is chaotic: small differences in starting conditions can produce large differences in outcome. Edward Lorenz, working in Cambridge ↗, Massachusetts, in the 1960s, was the first to describe this quantitatively, and the discovery made deterministic long-range forecasting provably limited rather than merely difficult.
Running the model many times from slightly different starting points converts a single answer into a spread, which is the honest output.
The ensemble approach accepts that limitation and turns it into signal. Instead of running the model once from the best-guess initial state, a forecasting centre runs it many times — typically with dozens of members — each started from a slightly different initial state. The differences are constructed to represent the actual uncertainty in the observations: they are not random noise but structured perturbations designed to probe where the analysis is weakest. Run fifty of them forward and you have fifty plausible futures rather than one.
What the spread tells you
Where the members cluster tightly, the atmosphere is behaving predictably. Where they diverge, the situation is genuinely uncertain — and that divergence, the spread itself, is real information. A forecaster who sees fifty members agreeing on rain can speak with confidence; a forecaster who sees thirty showing rain and twenty showing nothing can tell you that the probability is around sixty percent, which is a more honest statement than "rain" or "no rain."

This probability is the honest output Lorenz's theory demanded. The European Centre for Medium-Range Weather Forecasts, based in Reading, England, began running ensemble forecasts operationally in 1992, and the National Centers for Environmental Prediction in the United States followed. Both systems have grown steadily in the number of members and the resolution at which each member runs. ECMWF's ensemble is now one of the most closely watched products in operational meteorology worldwide.
Key concepts in the method
- Deterministic forecastone model run from one initial state; precise but overconfident
- Initial perturbationsstructured small differences seeded into each ensemble member to represent observational uncertainty
- Ensemble spreadthe range of outcomes across all members; wide spread = genuine uncertainty, narrow spread = predictable situation
- Stochastic physicsrandom variation applied inside parameterisation schemes to represent model uncertainty within each run
- Multi-model ensemblemembers drawn from different forecast models, capturing structural disagreements between them
- Calibrationwhether a forecast of X% probability actually verifies at X% over many cases; the target for ensemble reliability
The perturbations are only half the ensemble's machinery. Forecast models also carry their own errors — approximations in how they represent clouds, turbulence, or precipitation, the territory covered by parameterisation. A single model run with perfect initial data would still drift from reality, because the model itself is imperfect. Modern ensembles address this with stochastic physics: small random variations are introduced into the parameterisation schemes during each run, so the spread reflects model uncertainty as well as observational uncertainty. Some operational centres go further by combining members from different models entirely — a multi-model ensemble that captures structural disagreements between competing formulations of the atmosphere.
What a probability forecast actually means
A probability from an ensemble is a frequency statement. If the ensemble says forty percent chance of rain and you collect every forecast that ever said forty percent, roughly forty percent of those occasions should have produced rain. Checking whether that is true is the job of verification, and it turns out to be painstaking: an ensemble can be sharp (confidently giving high or low probabilities) while still being wrong, or it can be reliable (well-calibrated against observations) while telling you almost nothing useful. The World Meteorological Organization ↗ publishes guidelines on how ensemble prediction systems should be evaluated, because a badly calibrated ensemble is worse than a simple climatological estimate.
Chronology
- 1960sLorenz quantifies sensitive dependence on initial conditions in Cambridge, Massachusetts
- 1992ECMWF begins operational ensemble forecasting in Reading, England
- Ongoingensemble member counts and per-member resolution both increase as computing power grows
For the person checking a forecast before the weekend, the ensemble surfaces most visibly as a percentage — that italicised figure beside the rain icon — or as a temperature range rather than a single number. Both are compressions of a richer picture: fifty model runs, carefully spread across the range of plausible initial states, each contributing its vote. The single-number forecast never went away, but it is now understood as a summary of the distribution rather than a statement of fact. That shift — from one answer to a spread with a shape — is the most intellectually honest thing numerical weather prediction has done since it began.
Elsewhere in Where it fails
And why the failure has a shape. Everything in this sectionexplains how forecast error grows and where it concentrates.
- Where it fails, and why the failure has a shapeLongForecast error is not random: it grows fastest where the atmosphere is least stable, which is why some situations are predictable for days and others for hours.
- The butterfly, in practiceMediumSensitivity to initial conditions is a measured property with a timescale attached, not a metaphor.
- The Grid, and What Falls BetweenLongA model divides the atmosphere into boxes, and anything smaller than a box — a shower, a hill — has to be represented rather than resolved.
- VerificationShortA forecast is only as good as the record of how it did, and the scoring is its own small discipline.