Every point is one calendar month, scored exactly like the live 30-day window
but over that month's hourly forecast/observation pairs, bucketed by the time the
forecast was valid. Only complete months are shown (never the current one).
- Temp / Wind show RMSE (root-mean-square error) in real units —
lower is better. Temperature errors are naturally larger in winter, so the
bold line is a trailing 12-month average that cancels the seasonal cycle;
the faint dots are the raw monthly values. Switch View to see the monthly
numbers directly.
- Combined and Rain use 0–100 skill scores (higher is better)
— Combined is the lead-weighted headline, Rain is the F1 of rain/no-rain calls.
Rain has gaps in dry months, where there is no rain signal to score.
- Lead picks the forecast horizon: D-1 is "what it said 24 h ahead",
up to D-7. All combines horizons (1/d-weighted).
- View → Yearly overlays one model's calendar years on a Jan–Dec axis
(pick the model in the control that appears). Brighter lines are more recent
years and the same months line up vertically, so "did the last year or two
slip?" — say, after an agency's budget cuts — reads directly off whether the
bright lines ride above the dim stack (for errors; below, for skill). The
headline compares the latest year to the same calendar months of all earlier
years, so a partial year is never judged against other years' winters.
This page uses the error-based score, the live board does not. The
scoreboard's headline is now score v2, where temperature and wind are a tolerance
hit rate plus an extreme-event score rather than RMSE. That score defines "extreme"
against a location's own climate, and a single month can only measure that from
itself, so the yardstick would drift with the season and a five-year line would be
partly a chart of its own thresholds. A fixed RMSE cap is the same ruler in 2021 as
in 2026, which is what a trend needs — so every month here, old and new, is the
error score, and the numbers are not comparable with the live board's.
Two honest caveats. The best_match truth series' own composition can
shift over the years, so a trend partly reflects the yardstick, not only the model.
And these are "seamless" products — a rising or falling line includes the model
upgrades that happened over this period, which is exactly what a
"has it changed?" question is asking. HRRR is omitted because it duplicates GFS at
short lead (the seamless GFS product already uses it there) and has no longer-lead
archive.
Data from Open-Meteo's
previous-runs and
historical-forecast
APIs, baked offline. Covers every default scoreboard city with its full model roster;
each city's models reach back as far as that model's archive allows (GFS to 2021,
most others from 2024).