Executive Summary
On 3 September, Google DeepMind released WeatherNext 3, its most advanced global weather AI model to date. Its defining feature is a fresh forecast every hour, compared to the four daily updates of traditional physics-based models. We put it to the test against ECMWF's IFS, the industry standard for medium-range forecasting, on the three variables that matter most for a renewable portfolio: wind at 100m, solar irradiance, and temperature at 2m. WeatherNext 3 came out ahead on all three, at every horizon and from every initialization. In this post, we walk through how we ran the comparison, where the gains come from, and what they are worth in imbalance cost for a 100 MW wind or solar asset.
Power markets are organized sequentially, with trading stages operating over progressively shorter time frames—from the day-ahead market to intraday trading and, ultimately, real-time balancing. As delivery approaches, improved forecasts of generation and consumption allow market participants to update their positions, helping to minimize imbalances on the power grid.
Each market participant is responsible for managing its own imbalance, the deviation between its market and physical position. The participant establishes its commercial position in the day-ahead market and can subsequently adjust it through intraday trading. Closer to real time, it can also influence its physical position by steering generation, consumption, or curtailment. The system operator settles any remaining imbalance at the imbalance price, which can result in high costs.
Weather models therefore play an important role in anticipating changes in renewable generation and electricity demand. More accurate forecasts allow market participants to adjust their commercial or physical positions before delivery, reducing the volume exposed to imbalance settlement. Assessing a weather model’s value therefore requires examining whether its improved forecasts lead to lower imbalance volumes and, ultimately, lower imbalance costs.
On the 3rd of September, Google DeepMind released WeatherNext 3 (WN3), their most advanced global weather AI model yet. Like its predecessors, it is trained on and initialized from analysis data: a gridded best estimate of the current atmosphere, assembled by a traditional numerical weather prediction system from millions of observations. WN3 initializes a fresh analysis every six hours.
WN3 sets itself apart by using geostationary satellite imagery alongside analysis data. Each run processes a rolling window of satellite mosaics updated hourly. This allows WN3 to incorporate recent atmospheric developments and issue a new forecast every hour, rather than waiting for the next six-hourly analysis.
Before looking at the results, it helps to distinguish some timing concepts:
Every WN3 forecast starts from atmospheric fields produced by IFS. At 00, 06, 12 and 18 UTC—the four synoptic initialization times—the latest field is a new IFS analysis. For the hourly initializations in between, WN3 uses the corresponding forecast step from the most recent IFS run. A WN3 forecast initialized at 08 UTC, for example, starts from the two-hour forecast produced by the 06 UTC IFS run.
WN3 therefore builds on IFS rather than replacing it. IFS advances the atmospheric state using a numerical model of the underlying physics. WN3 supplements the IFS-derived state with geostationary satellite imagery extending five hours beyond it, then produces a forecast using a learned model. WN3 may outperform the corresponding IFS forecasts in the results presented here, but it still depends on IFS for its atmospheric starting state.
This use of newer satellite imagery affects forecast availability. For its latency-adjusted evaluation, the WN3 paper assumes six hours for the required inputs to become available and up to another hour for execution and dissemination. A WN3 forecast therefore reaches users roughly seven hours after its nominal initialization time, compared with about six hours for IFS. In the results below, the improvement in forecast quality more than compensates for this additional delay.
In the analysis, the four daily runs starting from a new IFS analysis are treated separately from the other 20 hourly updates. These interim updates use the +1 to +5-hour forecast steps from the latest IFS run, supplemented with more recent satellite imagery. Comparing the two groups helps distinguish the value of a new atmospheric analysis from the value of fresher satellite observations.
We evaluated both models against ERA5, ECMWF’s global reanalysis dataset. ERA5 combines observations with a numerical weather model to produce a consistent, gridded reconstruction of past atmospheric conditions. We used it as the reference at 0.25° resolution across 609 grid cells covering the Benelux: 432 over land and 177 over the North Sea, where forecast accuracy matters most for offshore wind. The evaluation covers 1 February to 1 August 2026 and lead times from 1 to 48 hours. Errors are measured using mean absolute error, expressed in each variable’s own units.
The three variables correspond directly to weather-dependent energy assets and demand:
The benchmark is ECMWF’s Integrated Forecasting System (IFS) control forecast, the unperturbed ensemble member. We chose IFS rather than ECMWF’s own AI model because European market participants still use IFS widely for operational scheduling. The question is therefore whether switching to WN3 would improve current practice, rather than which model performs best in a like-for-like comparison of forecasting architectures.
Let’s start with some raw values rather than errors. For one week in July, at one grid cell near Antwerp, we plotted ERA5 in black with both forecasts over it:
The difference is clearest for wind speed. Both models capture the overall pattern, but IFS underestimates wind speeds during most of the week, while WN3 follows the ERA5 values more closely. For GHI, both models reproduce the daily solar cycle reasonably well, but at this 24-hour lead time, WN3 does not benefit from its satellite imagery. For temperature, it’s hard to conclude, although WN3 seems to have a small edge.
Now that you know what the data looks like, we can evaluate the entire window:
| Variable | WN3 MAE | IFS MAE | WN3 better by |
| 100 m wind | 0.63 m/s | 0.89 m/s |
29% |
| GHI | 30.8 W/m² | 42.9 W/m² | 28% |
| 2m temperature | 0.61 °C | 0.74 °C | 17% |
GHI is evaluated only when ERA5 reports more than 20 W/m² of irradiance. Periods with little or no sunlight are relatively easy to predict and would make both models appear more accurate without revealing much about their ability to forecast solar generation.
WN3 has a lower mean absolute error for all three variables, with improvements ranging from 17% for temperature to 29% for wind speed.
However, the averages do not show how forecast accuracy changes with lead time. This matters because different market decisions rely on different forecast horizons. For day-ahead scheduling, for example, excellent performance over the next three hours is of limited value if accuracy deteriorates sharply by the following day. Figure 3 therefore compares the errors of both models as the forecast horizon increases.
Fig. 3: WN3 against ECMWF across three variables, grouped by the four synoptic initializations.
WN3 has a lower error than IFS in every panel, across all lead times and initialization types. The difference is striking: a WN3 wind forecast with a lead time of 48 hours is more accurate than an IFS forecast with a lead time of just three hours. In this case, switching models provides more than 45 hours of effective lead-time advantage.
What the hourly updates add
Next, we isolate the value of WN3’s headline feature: hourly updates that incorporate new satellite imagery between successive IFS analyses.
For each target hour, we followed the sequence of available forecasts as delivery approached, from a lead time of 31 hours down to eight hours. This broadly reflects the window used for day-ahead nominations and early intraday corrections. Over this period, WN3’s wind error decreases by 13%, its irradiance error by 8% and its temperature error by 9%.
Fig 4: Counting down to a single target hour, 12 UTC.
To determine where these improvements come from, we separate the changes occurring at the four synoptic initializations, when a new IFS analysis becomes available, from those occurring during the 20 interim hourly updates:
Wind accuracy improves mainly at the four synoptic initializations, while most gains in irradiance and temperature occur during the hourly updates in between.
A likely explanation lies in what satellites observe most directly. Geostationary imagery continuously tracks cloud location and development, which strongly affect solar irradiance and, through their effect on incoming radiation, near-surface temperature. It provides less direct information about wind speed at 100 m, which depends more on the three-dimensional atmospheric state and large-scale flow captured by IFS. A new analysis therefore provides a much stronger wind update than fresh satellite imagery alone.
This distinction matters most for intraday trading. Intraday corrections are made after the day-ahead nomination, using the most recent forecast available at the time. That is precisely where WN3’s hourly updates can add value. In this evaluation, the updates provide a meaningful advantage for managing a solar PV position, but contribute relatively little to managing a wind position.
To estimate the financial impact of the forecast improvements, we simulate a simplified day-ahead strategy. A renewable generator submits its day-ahead nomination following a generation forecast for its assets using the latest weather forecast available from each model. The nominated production is then compared with the actual production computed using the ERA5 weather data, which serves as the ground truth, and the resulting difference is settled at the imbalance price. The simulation assumes that the nomination is not adjusted through intraday trading.
Renewable forecast errors are often correlated across portfolios exposed to the same weather. Overproduction may coincide with a long system and lower—or even negative—imbalance prices, while underproduction may coincide with a short system and high prices. Deviations can therefore be particularly costly when they reinforce the system imbalance, although this will not be the case in every settlement period.
Informing day-ahead nominations with WN3 forecasts could therefore have a significant financial impact compared to using ECMWF’s IFS forecasts. In this analysis, we examine the financial impact of following nominations informed by each model across two case studies. The two cases at hand are an offshore wind farm and a solar installation at exemplary locations in Belgium. The analysis uses price data from February 2026 to July 2026 to capture varying weather conditions.
Wind — 100 MW offshore
The MAE on forecasted wind improves by 0.32 m/s when using WN3 instead of ECMWF. A wind farm climbs from zero to rated output between roughly 3.5 and 12 m/s and sits in that steep band around 55–60% of hours, putting its average sensitivity near 9 MW per m/s per 100 MW installed. That leaves 2.763 MW of average absolute error avoided in every hour, or about 12,000 MWh over 6 months.
≈ €465k avoided imbalance costs over 6 months, or €4,650 per MW installed.
Solar — 100 MWp
GHI MAE improves by 8.2 W/m². PV output tracks irradiance close to linearly, so at a performance ratio of 0.85, that is 0.7 MW of error avoided in each of roughly 4,400 daylight hours, or about 3,000 MWh over 6 months.
≈ €150k avoided imbalance costs over 6 months, or €1,500 per MWp.
However, note that a large improvement in the irradiance forecast comes from hourly updates, which would have a larger impact on intraday position than on day-ahead position. On the other hand, intraday corrections and curtailment would generally reduce the volume reaching imbalance settlement.
| MAE (MW) | MWh misscheduled (MWh) | MWh misscheduled (MWh) | Cost per MWh misscheduled (€/MWh) | ||
| Wind | WN3 | 7.237 | 31,272.0 | 1,511,567.0 | 48.34 |
| ECMWF | 10.0 | 43,212.0 | 1,976,589.0 | 45.74 | |
| Solar | WN3 | 1.810 | 7,822.0 | 130,530 | 16.69 |
| ECMWF | 2.508 | 10,838.0 | 282,411.0 | 26.06 |
Three things we would take from this.
WN3 is a meaningful improvement over current practice. Around 28–29% lower error on the two variables that drive renewable output is worth more than a day and a half of lead time: a two-day-old WN3 wind forecast still beats a three-hour-old IFS one. WN3 reaching the desk an hour later barely dents that.
The extra hourly updates have a major impact on short-term irradiance forecasting but less so for longer-term or wind forecasts. In practice, that means WeatherNext3 could significantly improve intraday solar corrections.
The accuracy is worth real money, and we have only measured part of it. Nominating a 100 MW wind farm on WN3 instead of ECMWF reduces imbalance costs by €465k over the first 6 months of 2026. For a 100 MW PV farm, the imbalance reduction is worth €150k. While these numbers are based on a simplified simulation, they provide a good indication of the value of improved day-ahead forecasts.
We will analyze the additional value unlocked by WeatherNext 3 in follow-up blog posts, where we will look at the value of intraday solar updates and assess whether WeatherNext 3 contains “unpriced information” that would allow counter-trend trading.
Interested in additional insights? Let us know which analysis interests you most, or contact Robbe Sneyders to evaluate WeatherNext 3 for your specific assets or portfolio.