suchablogArithmeticFitzRoyInstrumentsThe roomWhere it failsArchiveAbout

Where it fails, and why the failure has a shape

Forecast error is not random: it grows fastest where the atmosphere is least stable, which is why some situations are predictable for days and others for hours.

This piece
SectionWhere it fails
LengthLong
Figures3

Forecast errors are not scattered evenly across the map. They cluster, they grow in patterns, and those patterns reveal something true about the atmosphere.

The error is not noise

Pick any bad forecast — the sunshine that turned to hail, the wind that refused to arrive — and the temptation is to call it a miss and move on. That instinct is wrong, and productive forecasting science has spent decades learning to resist it. Errors are not random. They have a preferred geography, a preferred season, and a preferred type of weather. Mapping where the models fail is how you discover where the atmosphere is hardest to constrain.

Where it fails, and why the failure has a shape
Error is not evenly distributed. It grows fastest where the atmosphere is least stable, which is a property of the day.Photograph · suchablog asset kit

The most basic version of the argument goes back to Edward Lorenz, working at the Massachusetts Institute of Technology in the early 1960s. Running a primitive numerical model, he found that two runs started from initial conditions that differed only at decimal points diverged, over time, into completely different solutions. The word "chaos" came later; what Lorenz had found was that the atmosphere's equations are nonlinear, meaning that small differences compound rather than cancel. That discovery set the practical upper bound on deterministic forecasting ↗, but it did not say the bound was the same everywhere or for every kind of weather.

Instability amplifies, stability damps

The key variable is atmospheric instability. Where the atmosphere is stable — a dome of settled high pressure, cold air sitting quietly under a temperature inversion — small errors stay small. The correct and incorrect versions of the atmosphere drift apart slowly. A good model can see five, six, even seven days ahead and still be trusted on the broad pattern.

Where the atmosphere is unstable, the opposite is true. Warm, moist air forced upward can organize into thunderstorms within an hour; whether it does, and exactly where, depends on details of temperature and moisture at scales too fine for any operational model to resolve directly. A model might correctly predict that conditions are ripe for convection across a region the size of a small country, and still have no reliable opinion on which valley floods and which stays dry. The error is not in the physics; it is in the resolution. Anything smaller than a model's grid box — the smallest horizontal scale a model explicitly calculates — must be handled by parameterisation, and parameterisation is where approximation error is baked into the forecast before it even starts.

A hand-annotated surface chart
Contours are closed by hand where the model left them ambiguous. The pencil line is a judgement, not a tracing.Photograph · suchablog asset kit

This is why forecast skill degrades faster in the tropics and over convective continental interiors than it does over the cool, stable oceans of the mid-latitudes in winter. The jet stream, though energetic, is at least large-scale; thunderstorm clusters are not.

The forecast problem is really an initial-condition problem

The other major shape of failure is positional. A synoptic-scale system — a low-pressure area, a front — will be in the right place in the model on day one, a little displaced on day three, and on day five it may be an entire weather regime out of phase with reality. Not because the dynamics are wrong, but because data assimilation ↗ — the process of reconciling millions of observations with the model's prior state — cannot make the initial state perfect. Every observation has an error. Every observation is taken at a point, not in a continuous field. The model's first-guess field, the background state from which assimilation corrects, is itself the product of a previous forecast that already carried some small error. These inherited imperfections, amplified by the nonlinearity Lorenz identified, manifest most visibly at the synoptic scale as positional errors: the storm track that runs slightly too far north, the front that arrives half a day late.

Forecasters know the classic failure modes. Explosive cyclogenesis — a low that deepens far faster than predicted — is a recurring one. The Bergen school, founded by Vilhelm Bjerknes in Norway in the early twentieth century, first gave meteorologists a way to identify the conditions that favour rapid deepening: the interplay of upper-level divergence, surface baroclinicity, and latent heat release. Models now simulate all of this, yet cyclones still occasionally surprise them because each of those ingredients is itself uncertain at initialization. When the errors in three compounding factors happen to reinforce, the answer can be badly wrong.

The ensemble as a map of failure

The practical response to all of this was the ensemble, now standard at major forecasting centres including the European Centre for Medium-Range Weather Forecasts in Reading, England, and the US National Centers for Environmental Prediction. Rather than running one forecast from one initial state, an ensemble runs dozens of slightly perturbed versions simultaneously. The spread of those runs is not a failure of nerve; it is an honest statement about which situations are predictable and which are not. A tight cluster of ensemble members at day five means the atmosphere is, in that situation, relatively constrained. A scattered fan of possibilities means it is not.

A radiosonde balloon at launch
The package weighs a few hundred grams and is used once. Its drift on the way up is how the wind aloft gets measured.Photograph · suchablog asset kit

The spread is, in effect, a real-time map of where the model expects to fail. Forecasters have learned to read it that way. When the ensemble disagrees violently about the track of an approaching system, the right public message is not the average of the members but an honest description of the range of outcomes. That is a harder message to communicate than a single number, and the discipline of communicating it well is still developing.

Verification as learning

None of this improvement would have happened without verification — the formal, statistical accounting of what the forecast said and what the atmosphere did. Robert FitzRoy, who established the first governmental storm-warning service in Britain in the 1860s, kept his own informal tally. Modern verification is industrialised: the World Meteorological Organization coordinates standards that let agencies compare performance across models, lead times, regions, and weather types. Verification scores for global NWP models are published and scrutinised, because without them there is no signal against which to improve.

From the working notes

What makes a forecast hard — and why

  1. Nonlinearitysmall initial errors compound over time rather than cancelling; first demonstrated formally by Edward Lorenz in the early 1960s
  2. Instabilityunstable air (convection, thunderstorms) amplifies errors faster than stable regimes (high pressure, cold inversions)
  3. Resolution limitmodel grid boxes can't capture features smaller than their own width; those features are parameterised instead
  4. Positional drifta correctly-shaped system placed slightly wrong at day one will be increasingly out of place by day five
  5. Explosive cyclogenesisrapid deepening of lows is a classic failure mode because three compounding uncertain factors must all be right at once

What verification has confirmed is the structural character of forecast error. It is not that models fail at random. They fail in predictable situations — rapidly deepening lows, convective outbreaks, sharp frontal passages — and that predictability means failure itself has a shape. Knowing the shape is not a consolation. It is the beginning of the fix: finer resolution where it matters, better parameterisation of the processes that can't be resolved, more observations in the data-sparse regions where the initial state is least well known, and ensembles wide enough to express the true range of what might happen next.

From the working notes

The machinery that manages failure

  1. Ensemble forecastingmany runs from slightly different starting states; spread = real-time estimate of predictability
  2. Data assimilationreconciling observations with the model's prior state; imperfect because every observation has error and gaps
  3. Verificationformal scoring of past forecasts against observations; the only way to know whether changes to a model represent improvement

The forecast you look at is the result of all of that machinery working in sequence. When it is wrong, it is usually wrong in a way the forecaster already suspected.

Elsewhere in Where it fails

And why the failure has a shape. Everything in this sectiondescribes a limit that forecasters already know about.

  • The butterfly, in practiceMediumSensitivity to initial conditions is a measured property with a timescale attached, not a metaphor.
  • The Grid, and What Falls BetweenLongA model divides the atmosphere into boxes, and anything smaller than a box — a shower, a hill — has to be represented rather than resolved.
  • EnsemblesMediumRunning the model many times from slightly different starting points converts a single answer into a spread, which is the honest output.
  • VerificationShortA forecast is only as good as the record of how it did, and the scoring is its own small discipline.
Where each other section starts