Annex III: Weather Year Selection Methodology //

TYNDP 2026 introduces a climate‑adaptive weather year selection methodology that integrates future climate projections (CMIP6 runs with SSP2‑4.5) alongside historical reanalysis weather information. This enables energy system modelling that better captures expected changes in renewable generation conditions and temperature‑driven electricity demand.

Input Data

The methodology is based on weather and energy variables from the PECDv4.2 dataset (Annex II), featuring the following models:

  • CMCC-CM2-SR5 (CMR5) by the Centro Euro-Mediterraneo sui Cambiamenti Climatici, Italy,
  • EC-Earth3 (ECE3) by the European community Earth System Model, and
  • MPI-ESM1-2-HR (MEHR) by the Max Planck Institute for Meteorology, Germany

The methodology uses:

  • Hourly PV, onshore wind, offshore wind power capacity factors
  • Weekly hydro reservoir and run‑of‑river generation profiles
  • Daily temperatures, converted into Heating Degree Days (HDD) and Cooling Degree Days (CDD), using regional thresholds (e. g., Austria: 12.3 °C heating, 25.5 °C cooling)
  • Installed RES capacities (PEMMDB) and load projections (ETM) for each target year

Candidate Climate Years

For each study horizon (2030, 2035, 2040, 2050), a window of 10 years per climate model is used (e. g., TY2030: 2025–2034).

Together with the three climate models (CMR5, ECE3, MEHR) , this yields 30 candidate climate‑year series per target year, from which three representative years are selected.

High level step‑by‑step selection method

1// Computation of annual averages and cumulative anomalies

  • For each climate year
  • For each variable (PV, wind offshore, wind onshore, hydro power run of river generation, hydro power reservoir generation, HDD+CDD)

2// Calculation of overall statistics (mean, standard deviation) and deltas

  • Deltas quantify the deviation from the full candidate range

3// Application of regional weighting

  • Based on installed RES capacities and load relevance
  • Only ENTSO‑E regions included (exclusions: Turkey, Ukraine, Moldova)

4// Normalisation of all parameters

  • Subtraction of the global mean and division by standard deviation to ensure comparability across variables with different units and magnitudes.

5// Dimensionality reduction via Principal Component Analysis (PCA)

  • Climate variables are reduced to key principal components, capturing the dominant patterns in RES availability and temperature‑driven load.

6// Clustering using k‑means methodology (k = 3)

  • Climate years are grouped in PCA space based on the similarity of normalised, weighted characteristics.
  • For each cluster, the medoid (year closest to the ­centroid) is selected.

Outcome

The method provides three representative weather years per TYNDP target year, each reflecting a distinct climate pattern (see Table 33).

Although the CMR5 model appears over‑represented in final selections, this has no climatic significance and does not bias the methodology.

Code nameSSP scenarioClimate modelClimate yearStudy Target Year
WS003SSP245CMR520272030
WS021SSP245MEHR20252030
WS029SSP245MEHR20332030
WS032SSP245CMR520312035
WS037SSP245CMR520362035
WS059SSP245MEHR20382035
WS065SSP245CMR520392040
WS071SSP245ECE320352040
WS077SSP245ECE320412040
WS091SSP245CMR520452050
WS092SSP245CMR520462050
WS106SSP245ECE320502050

Table 33: Climate year codes

Motivation and Role in Scenario Building

Energy system planning studies must account for weather-driven variability in both demand (temperature-dependent heating and cooling) and supply (wind, solar, and hydro generation). Simulating the 30 scenarios from the three climate models for each target year is computationally heavy and time consuming. The Weather Scenario Selection methodology reduces this set to k = 3 representative weather years per target year, each carrying a probability weight that reflects the frequency of its weather characteristics. This reduction aims make the computational problem tractable while keeping an appropriate level of quality on the final outcomes.

The selected scenarios and their weights allow downstream market modelling processes to approximate expected values over weather uncertainty. This reduces the simulation burden from 30 to 3 runs per target year.

The Weather Scenario Selection is performed early in the scenario building process. Its outputs are the three representative weather scenarios and their associated weights are used as inputs to:

  • Demand profiling: hourly electricity, hydrogen and thermal demand time series for hybrid heating are constructed for each selected weather scenario, reflecting temperature-driven heating and cooling loads;
  • RES generation profiles: wind, solar, and hydro generation time series are extracted by linking PECDv4.2 with PEMMDB for the selected weather scenarios;
  • Market modelling: adequacy and market simulations are run for each of the three representative weather scenarios, and results are combined using the probability weights.

Scope of Analysis

The scope of analysis for the weather year methodology is further elaborated in Table 34 below.

ParameterValue
Target years2030, 2035, 2040, 2050
Climate year windows2030 (2025–2034), 2035 (2030–2039), 2040 (2035–2044), 2050 (2045–2054)
Climate modelsCMCC-CM2-SR5 (CMR5), EC-Earth3 (ECE3), MPI-ESM1-2-HR (MEHR)
Candidates per target year3 models × 10 years = 30 weather scenarios
Number of representative years (k)3
Bidding zones53 European zones
Technologies / Parameters analysed6: HDD / CDD, Run-of-River, Reservoir, Wind Offshore, Solar PV, Wind Onshore
Macro regions9 (Baltic, CE, CW1, CW2, Nordic, SC, SE, SW, UK)

Table 34: Scope of analysis for weather year methodology

Methodology Overview

The selection follows a five-stage pipeline: (1) Data Ingestion consolidating inputs from the PECD, demand and generation time series, and other TSO-provided data; (2) Statistical Processing including regional aggregation, weighting, and z-score normalisation; (3) Dimensionality Reduction via PCA; (4) k-means Clustering of the 30 weather scenarios into 3 groups; and (5) Representative Selection of the medoid from each cluster together with probability weights, shown in Figure 19 above.

Figure 19: Overview of the weather year selection methodology

Stage 1: Data Ingestion

Three categories of weather-dependent data are loaded for each of the 30 candidate weather scenarios and 53 bidding zones:

Heating and Cooling Degree Days (HDD / CDD)

Hourly temperature data from the three climate models is converted to daily averages. A threshold transformation is then applied using zone-specific heating and cooling thresholds derived from econometric calibration (per-country GDP, building thermal transmittance, and regression coefficients). Temperatures above the cooling threshold produce cooling degrees for the definition of CDD; temperatures below the heating threshold produce heating degrees for the definition of HDD; temperatures within the comfort band produce zero contribution to both HDD and CDD.

Renewable Generation (Solar PV, Wind Onshore, Wind Offshore)

PECD capacity factor profiles are multiplied by installed generation capacity (MW) from the TYNDP generation capacity dataset for the target year. Sub-types are aggregated to technology level (e. g., Solar PV utility tracking, non-tracking, industrial rooftop, and residential rooftop are summed into a single “Solar” category; fixed and floating offshore wind are summed into the “Offshore” category).

Hydro Inflows (Run-of-River, Reservoir)

Inflow time series are loaded directly from PECD files for each zone, validated against installed hydro capacity in the generation dataset.

Stage 2: Statistical Processing

For each of the six technologies, two indicators are computed per weather scenario to characterise its mean level and its intra-annual variability relative to the climatological baseline.

Macroregional Aggregation

The 53 bidding zones are grouped into nine macroregions to reduce dimensionality while preserving geographic diversity (Table 35):

MacroregionCountries
BalticLatvia, Estonia, Lithuania
CECzech Republic, Slovakia, Hungary, Poland, Romania
CW1Belgium, France, Netherlands
CW2Germany, Austria, Switzerland, Luxembourg
NordicDenmark, Finland, Norway, Sweden
SCItaly
SEGreece, Cyprus, Bulgaria, N. Macedonia, Croatia, Slovenia, Albania, Bosnia-Herzegovina
SWSpain, Portugal
UKUnited Kingdom, Ireland

Table 35: Macroregional aggregation

For each macroregion, the regional time series is the sum across all constituent bidding zones.

Mean-Anomaly Delta

For each region “r” and weather scenario “s”, the annual mean is computed and compared to the “grandmean” across all scenarios. The mean-anomaly delta is the difference between a scenario’s annual mean and the average across all 30 scenarios. This captures whether a given weather year produces above- or below-average generation (or demand) for a given technology in a given region.

Variability-Anomaly Delta

The variability indicator captures how much the intra-annual variability of a scenario deviates from the expected level. It is computed as the cumulative anomaly (the running sum of deviations from the “grand-mean” over all timesteps) minus the standard deviation of per-scenario standard deviations across all scenarios. This distinguishes between years with stable output and years with highly volatile output.

Capacity Weighting

Regional deltas are weighted to reflect the relative importance of each macroregion. The weighting approach differs by technology type:

  • For renewable and hydro technologies: the weight of a region equals its share of total installed capacity for that technology. Regions with more installed capacity exert a proportionally greater influence on the clustering.
  • For HDD / CDD: the weight of a region equals its share of total electricity demand (from the Bidding Zone Overview dataset), reflecting the fact that temperature sensitivity scales with consumption volume.

The weighted deltas collapse the regional dimension into a single value per scenario for each technology.

Z-Score Normalisation

Each weighted delta series is standardised to zero mean and unit variance. This ensures that all six technologies contribute equally to the subsequent dimensionality reduction, regardless of their original physical units and scales.

Stage 3: Dimensionality Reduction (PCA)

After the previous Stage 2, each of the 30 weather scenarios is described by a 12-dimensional feature vector comprising six normalised mean-anomaly values and six normalised variability-anomaly values, one per technology.

These 12 dimensions are organised into two independent subspaces: (i) the mean-anomaly matrix A (30 scenarios × 6 technologies), whose columns are the z-scored mean deltas for each technology; and (ii) the variability-anomaly matrix S (30 scenarios × 6 technologies), whose columns are the z-scored variability deltas.

A Principal Component Analysis (PCA) is applied independently to each subspace, retaining only the first principal component (PC1). Each weather scenario is then represented as a point in a two-dimensional space: the first axis (PC1 of the mean-anomaly matrix) captures the dominant mode of co-variation in mean generation / demand levels across technologies; the second axis (PC1 of the variability-anomaly matrix) captures the dominant mode of co-variation in intra-annual variability.

Stage 4: K-Means Clustering

The 30 points in the reduced 2D space are partitioned into k = 3 clusters by k-means clustering, minimising the within-cluster sum of squares. The algorithm is run with 10 random initialisations to mitigate sensitivity to initial centroid placement, and the partition with the lowest total within-cluster variance is retained. A fixed random seed ensures full reproducibility.

The representative year for each cluster is the scenario closest to the cluster mean in the two-dimensional PCA space. By selecting an actual weather scenario rather than the centroid itself (which is a synthetic point with no corresponding time series), the representative weather scenario chosen always corresponds to an existing model year with a complete, physically consistent set of profiles across all technologies and zones.

Stage 5: Probability Weight Computation

The probability weight of each representative scenario equals the proportion of historical weather years assigned to its cluster: w_k = \lvert C_k \rvert / N, where \lvert C_k \rvert is the number of years in cluster k and N = 30 is the total number of candidate scenarios.

This is a non-parametric density estimator under the ergodicity assumption: the temporal frequency of weather patterns in a sufficiently long temporal record approximates their probability of occurrence. The weights are non-negative, sum to unity by construction, and represent the empirical likelihood of each weather regime — not a measure of weather severity or extremity.

Results

The following Table 36 summarises the selected representative weather years and their probability weights for each target year. For each target year, the k-means algorithm identifies three distinct weather regimes.

One cluster typically contains roughly half of the 30 candidates, representing the most common weather pattern, while the other two clusters capture less frequent but distinct regimes.

Target YearClusterRepresentative YearWeightCluster Size
20300WS0290.309
20301WS0030.6319
20302WS0210.072
20350WS0370.4313
20351WS0590.3310
20352WS0320.247
20400WS0770.4012
20401WS0710.4012
20402WS0650.206
20500WS1060.206
20501WS0910.237
20502WS0920.5717

Table 36: Representative Weather Scenarios for TYNDP 2026

Integration with the Scenario Building Process

The three representative weather years and their weights are integrated into the broader scenario building workflow as follows:

1// PECD data extraction: For each target year, the full hourly PECD time series (renewable capacity factors, hydro inflows, temperature profiles) are extracted for only the three selected weather years, rather than all 30 candidates.

2// Demand profile construction: The temperature time series of the selected weather years feed into the demand profiling process, where they drive the construction of weather-dependent electricity demand components — in particular heating (heat pumps, electric boilers, hybrid heat pumps) and cooling (air conditioning) loads. Each selected weather year produces a distinct annual demand profile reflecting its temperature characteristics.

3// RES and hydro profiling: Wind, solar, and hydro generation profiles are built from the PECD capacity factors of the three selected years, scaled by the installed capacities defined in the scenario.

4// Market simulations: Adequacy and market studies are run three times per target year, once for each representative weather scenario. The simulation results (e. g., generation dispatch, flows, marginal costs, adequacy margins) are then combined using the probability weights to produce expected values. This weighted-average approach allows planners to derive robust expected indicators while preserving the ability to examine individual weather regimes.

Key Assumptions and Design Choices

  • k = 3 clusters: The number of representative years is set to three, balancing computational cost against weather variability. This is consistent with previous TYNDP cycles and provides a manageable set for downstream modelling while still distinguishing between fundamentally different weather regimes.
  • Equal technology weighting via z-score normalisation: All six technologies receive equal influence in the PCA through standardisation.
  • Two independent PCA reductions: Separating the mean-anomaly and variability-anomaly subspaces into two independent PCA reductions preserves the interpretability of the resulting axes: one captures level effects (wet vs. dry, windy vs. calm years) and the other captures variability effects (stable vs. volatile years).
  • Ergodicity assumption: The 30 weather scenario series are assumed to constitute a representative sample of the climate distribution for each target year window. Under this assumption, cluster sizes directly estimate the probability of each weather regime.
  • Macroregional aggregation: Reducing 53 bidding zones to 9 macroregions prevents the analysis from being dominated by node-level noise while preserving the major geographic contrasts in European weather patterns (e. g., Nordic vs. Mediterranean, Atlantic vs. Continental).

Meteorology of the Selected Weather Years

The following graphs show the selected weather years, its weighted average according to Table 36 and the whole 30 candidate weather years for any target year.

Any weather year is characterized by 6 parameters grouped in 3 figures: offshore and onshore wind, HDD / CDD and solar irradiance, hydropower run-of river and reservoir generation.

Offshore Wind vs. Onshore Wind

Figure 20 shows the deviation of the 30 candidate weather years from their ensemble mean for each target year for offshore and onshore wind speed. PECD data has been aggregated separately for P2OF and P2ON zones to derive the combined deviations for all 27 EU countries.

The result of the selection is highlighted with blue, orange, and green circles, while the weighted selection is indicated by the black square. The weighted selection is striking for target year 2040, with a deviation of -0.16 m s-1 in offshore wind and –0.19 m s-1 in onshore wind.

Figure 20: Offshore and onshore wind in 2030, 2035, 2040 and 2050

HDD / CDD vs. Solar irradiance

Figure 21 shows the deviation of the 30 candidate weather years from their ensemble mean for each target year for global horizontal irradiance and HDD / CDD. For calculating the deviations for the global horizontal irradiance, P2ON regions have been aggregated for all EU-27 countries. For calculating the deviations for HDD / CDD, hourly air temperatures from PECD have been averaged to daily temperatures before deriving HDD and CDD. HDD and CDD have been aggregated by calculating the double sum over the specific climate year and all 27 EU countries.

The result of the selection is highlighted with blue, orange, and green circles, while the weighted selection is indicated by the black square. The weighted selection for target year 2030 indicates 1.78 W m-2 less solar irradiance than the ensemble mean, while the weighted selection for target year 2040 indicates 3626 °C x days less than the ensemble mean for heating and cooling across the 27 EU countries.

Figure 21: HDD / CDD vs solar irradiance in 2030, 2035, 2040 and 2050

Hydropower run-of river vs. reservoir

Figure 22 shows the deviation of the 30 candidate weather years from their ensemble mean for each target year for hydropower reservoir generation energy and hydropower run-of-river generation energy. Weekly PECD data of the produced amount of energy has been aggregated for SZON bidding zones to derive the combined deviations for all 27 EU countries. The result of the selection is highlighted with blue, orange, and green circles, while the weighted selection is indicated by the black square.

The weighted selection indicates a hydropower reservoir energy generation slightly above average for target years 2030, 2040 and 2050 and slightly below average for target year 2035. The hydropower run-of-river energy generation of the weighted selection is very close to the average for all target years with a maximum of 3.56 GWh deviation in 2035.

Figure 22: Hydropower run-of-river vs reservoir in 2030, 2035, 2040 and 2050