Accuracy
How well the forecast does, measured.
The rule in the methodology is that no forecast number appears on this site unless it can be checked. This page is the check: each horizon's typical error, measured over the last four weeks the same way operations forecast — and published inside the very file the forecast ships in.
Latest measurement: September 28, 2026 at 7:00 AM, refreshed twice a day.
- 24 hours ahead±2,537 MW
Forecast with real data only.
- Yardstick
- ±2,694 MW(6% better)
- Bias
- 0 MW
- Coverage
- 80%
- 48 hours ahead±3,001 MW
Recursive: resting on day 1's forecast.
- Yardstick
- ±3,761 MW(20% better)
- Bias
- -312 MW
- Coverage
- 80%
- 72 hours ahead±3,111 MW
Recursive: the model talks to itself for two days.
- Yardstick
- ±4,308 MW(28% better)
- Bias
- -280 MW
- Coverage
- 80%
The error is the MAE — the average distance, in MW, between what the model forecast and what ONS recorded, half-hour by half-hour. For scale: the system's mean curtailment in 2026 runs a few thousand MW, with peaks above 30 GW on the big days.
How the measurement works
The backtesting forecasts the way operations forecast. The model trains only on what existed up to thirty days ago; from there, it anchors on each day of the window and projects three days ahead — the first with real data only, the other two resting on its own forecasts, exactly how the published forecast is made. Each horizon's error is the average of those thirty trials, and the model is only live because it beat the hardest yardstick in a series with memory: repeating yesterday. That yardstick is measured on the same instants and published in the same file (mae_30d_regua_mw) — if the model ever stops beating it, this site is the first place it shows.
Two honest footnotes. The measurement runs on ONS's final series, which revises closed months for weeks — measurement on each day's frozen data begins with the pipeline's load history. And the 7, 30 and 90-day windows per source, with the backtesting file for download, arrive once the series of published forecasts builds its own history.
Forecast vs. realized, day by day
The 30-day measurement scores the model; this table shows the product. Each row is a day already closed: the curtailment the model forecast 24h ahead (the most recent run while the data was still the same) against what ONS recorded afterward — a series that grows one day at a time, starting 28 Aug 2026.
| Day | Forecast | Realized | MAE |
|---|---|---|---|
| Sat, Sep 26 | 7,505 MW | 0 MW | ±7,505 MW |
| Fri, Sep 25 | 6,975 MW | 5,924 MW | ±1,325 MW |
| Thu, Sep 24 | 3,655 MW | 5,045 MW | ±1,689 MW |
| Wed, Sep 23 | 6,911 MW | 2,199 MW | ±4,713 MW |
| Tue, Sep 22 | 8,420 MW | 5,924 MW | ±2,901 MW |
| Mon, Sep 21 | 8,026 MW | 9,720 MW | ±1,996 MW |
| Thu, Sep 17 | 5,836 MW | 11,213 MW | ±5,377 MW |
| Wed, Sep 16 | 7,447 MW | 8,113 MW | ±1,843 MW |
| Tue, Sep 15 | 8,457 MW | 8,068 MW | ±1,230 MW |
| Mon, Sep 14 | 5,969 MW | 8,617 MW | ±2,692 MW |
| Sun, Sep 13 | 8,040 MW | 11,515 MW | ±3,474 MW |
| Sat, Sep 12 | 5,780 MW | 9,130 MW | ±3,381 MW |
| Fri, Sep 11 | 7,717 MW | 7,703 MW | ±1,671 MW |
| Thu, Sep 10 | 8,921 MW | 7,697 MW | ±1,471 MW |
Forecast and Realized are whole-day averages; the MAE is not the difference between them — it's the average error across the day's 48 half-hours, which don't cancel out the way an average would. A day can have close daily averages and still miss a lot half-hour by half-hour (or the opposite). The chart shows the two averages side by side; the exact half-hour math is in the file.
Full series to download at historico-acuracia.csv (also as .json).
Bias and the uncertainty band
MAE says how big the error is; bias says which way it leans — positive is the model overestimating curtailment, negative is underestimating. A model can have a small MAE and a large bias if it always errs the same way; the two numbers travel together in the file (vies_30d_mw) so that distinction never hides behind an average.
The forecast also publishes a P10–P90 band per half-hour: two extra LightGBM models, tuned to forecast the quantile instead of the mean. Alone, they are not enough — measured on 31 Aug 2026, the raw band covered only 42–53% of realized values instead of the nominal 80%, because two independently fit quantiles don't guarantee joint coverage. The fix is conformal: on the same backtesting trials, it measures how far the realized value escaped the raw band, and the 80th percentile of that escape becomes the margin that widens P10 and P90 (margem_p80_mw) until coverage hits the promise — the number cobertura_80_pct audits every run.