suchablogArithmeticFitzRoyInstrumentsThe roomWhere it failsArchiveAbout

Verification

A forecast is only as good as the record of how it did, and the scoring is its own small discipline.

This piece
SectionWhere it fails
LengthShort
Figures2

A forecast only tells you something if you later check whether it was right

Every forecast is a promise about the future. Verification is the practice of going back — after the weather has happened — and comparing what was promised to what arrived. Without it, a forecasting system is just confident noise.

Calculator and pen resting on a printed spreadsheet with computer screens blurred in the background
A forecast is only as good as the record of how it did. Scoring is a small discipline of its own.Photograph · suchablog asset kit

The discipline has real structure. A simple yes/no question, such as whether it rained, is scored with tools like the Brier score or the hit rate, which measure not just whether the call was made but how firmly it was made and whether that confidence was warranted. A model that says "70% chance of rain" every single day, regardless of conditions, will score poorly because probability without discrimination is meaningless. The metrics punish hedging and reward genuine skill.

A forecast is only as good as the record of how it did, and the scoring is its own small discipline.

Skill, in the technical sense, means performance relative to a baseline — usually climatology or a simple persistence forecast that just repeats yesterday. A forecast that merely does as well as "expect what normally happens in October" has added nothing. The World Meteorological Organization ↗ coordinates global verification standards partly to make scores comparable across national services, since a number means little if every centre calculates it differently.

A hand-annotated surface chart
Contours are closed by hand where the model left them ambiguous. The pencil line is a judgement, not a tracing.Photograph · suchablog asset kit

At the ECMWF in Reading ↗, England, verification runs continuously: each forecast is logged, the scores accumulate, and the record of model performance over years is what guides decisions about when to upgrade the system and whether the upgrade was actually an improvement. A change that looks better in theory but scores worse in practice gets rolled back. The numbers govern.

From the working notes

How scoring works

  1. Brier scorea probability score for yes/no events; lower is better, zero is perfect
  2. Hit ratefraction of observed events that were forecast; blind to false alarms without a companion metric
  3. Skill scoreforecast performance measured relative to a simple baseline (climatology or persistence); zero means no improvement over the baseline
  4. Equitable scoringa class of metrics designed so random or constant forecasts earn zero skill

One subtlety undermines every score: observation error. Verification assumes the measurement you are comparing against is the ground truth. But a radiosonde reading has its own uncertainty, a rain gauge under a tree catches less than falls, and a model grid cell covers many square kilometres while the gauge occupies a point. A forecast can be physically reasonable and still score poorly because the verifying observation was itself imperfect.

From the working notes

Key distinction

  1. Verification vs. validationvalidation tests whether the model's physics are coherent internally; verification tests whether the finished forecast matched reality

This is why meteorologists speak of "equitable" scoring — methods designed so that a lucky random guess earns no credit and a systematic bias receives no partial marks for occasionally landing right. The goal is a number that honestly reflects whether the atmosphere was understood, not one that flatters the system that produced the forecast.

Elsewhere in Where it fails

And why the failure has a shape. Everything in this sectionconcerns the limits of forecast skill: where predictions break down and how that breakdown follows a pattern.

  • Where it fails, and why the failure has a shapeLongForecast error is not random: it grows fastest where the atmosphere is least stable, which is why some situations are predictable for days and others for hours.
  • The butterfly, in practiceMediumSensitivity to initial conditions is a measured property with a timescale attached, not a metaphor.
  • The Grid, and What Falls BetweenLongA model divides the atmosphere into boxes, and anything smaller than a box — a shower, a hill — has to be represented rather than resolved.
  • EnsemblesMediumRunning the model many times from slightly different starting points converts a single answer into a spread, which is the honest output.
Where each other section starts