<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">HESS</journal-id><journal-title-group>
    <journal-title>Hydrology and Earth System Sciences</journal-title>
    <abbrev-journal-title abbrev-type="publisher">HESS</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Hydrol. Earth Syst. Sci.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7938</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/hess-30-5067-2026</article-id><title-group><article-title>BiasCast: learning and adjusting real time biases from meteorological forecasts to enhance runoff predictions</article-title><alt-title>BiasCast</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Konold</surname><given-names>Oliver</given-names></name>
          <email>oliver.konold@boku.ac.at</email>
        <ext-link>https://orcid.org/0009-0000-2241-697X</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Feigl</surname><given-names>Moritz</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-3545-864X</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff3">
          <name><surname>Podest</surname><given-names>Patrick</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Klingler</surname><given-names>Christoph</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Schulz</surname><given-names>Karsten</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-6616-2876</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Institute of Hydrology and Water Management, BOKU University, Vienna, Austria</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>baseflow AI solutions, Vienna, Austria</institution>
        </aff>
        <aff id="aff3"><label>3</label><institution>ELLIS Unit, LIT AI Lab, Institute for Machine Learning, Johannes Kepler University (JKU), Linz, Austria</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Oliver Konold (oliver.konold@boku.ac.at)</corresp></author-notes><pub-date><day>12</day><month>August</month><year>2026</year></pub-date>
      
      <volume>30</volume>
      <issue>15</issue>
      <fpage>5067</fpage><lpage>5096</lpage>
      <history>
        <date date-type="received"><day>8</day><month>October</month><year>2025</year></date>
           <date date-type="rev-request"><day>27</day><month>November</month><year>2025</year></date>
           <date date-type="rev-recd"><day>4</day><month>July</month><year>2026</year></date>
           <date date-type="accepted"><day>2</day><month>August</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Oliver Konold et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026.html">This article is available from https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026.html</self-uri><self-uri xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026.pdf">The full text article is available as a PDF file from https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e130">The use of deep learning models in hydrology is becoming an ever more prevalent application in operational flood forecasting. Such operational systems face performance degradation when transitioning from high quality reanalysis to meteorological forecast data with lower accuracy. This study investigates training strategies and Long Short-Term Memory network architectures to mitigate meteorological forecast-induced bias in maximum daily discharge predictions using the Extended LamaH- CE dataset and a subset of 451 basins. We systematically evaluated cross-domain generalization, transfer learning approaches, Encoder–Decoder LSTMs, Sequential Forecast LSTMs, and the role of input embeddings and integrating past discharge observations. The results show that domain shifts between reanalysis and forecast data lead to substantial skill loss, with median Nash–Sutcliffe Efficiency decreasing from 0.58 to 0.33. Among the tested strategies, the Sequential Forecast LSTM demonstrated the most stable improvements, achieving a median NSE of 0.63. Integrating recent discharge observations further enhanced performance, raising median NSE to 0.71 and surpassing even the reanalysis-driven baseline. In contrast, integrating archived forecasts or using more complex input embeddings did not yield consistent benefits and in some cases degraded model stability. Basin-level analysis reveals that forecast skill improvements compared to our baseline are not uniformly distributed across catchment types: the largest gains are concentrated in arid and precipitation-limited catchments, while alpine and snow-dominated catchments, despite experiencing the largest meteorological domain shift, show smaller improvements. This is likely because the LSTM cell state retains strong seasonal signals and thereby compensates for forecast input bias through its long-term memory mechanism, particularly in alpine and snow-dominated catchments where strong seasonal cycles dominate the hydrological response. These findings highlight the value of training strategies that allow models to directly learn bias correction during forecast transitions, emphasize the operational potential of combining sequential processing with near real-time discharge observations and identify physiographic catchment characteristics as key modulators of forecast skill improvement across diverse hydroclimatic settings.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e142">Accurate runoff prediction stands as one of the most critical challenges in modern hydrology, with far-reaching implications for flood risk management, water resource planning, and the design of resilient hydraulic infrastructure (Beven, 2012; Guo et al., 2021; Tran et al., 2025). While recent advances in deep learning have demonstrated that Long Short-Term Memory Networks (LSTMs) can effectively integrate multiple meteorological datasets to improve runoff simulation accuracy by learning complex spatial and temporal patterns (Kratzert et al., 2021), a fundamental challenge remains: operational forecasting systems rely on biased meteorological forecasts rather than reanalysis or observational data. This dependency introduces a cascade of uncertainties, as meteorological forecasts inherently exhibit lower accuracy and higher uncertainty than observational or reanalysis datasets (Lavers et al., 2021), with forecast errors further amplifying as lead time increases (Nester et al., 2012). The consequences of these uncertainties are particularly severe in flood forecasting applications, where timely and magnitudinally correct runoff predictions are critical for early warning systems and risk management (Chen et al., 2016).</p>
      <p id="d2e145">The biases in meteorological forecasts stem from factors such as model resolution, data assimilation techniques, or orographic effects, and they differ depending on the numerical weather prediction model (e.g., ECMWF-HRES, DWD-ICON, NOAA-GFS), the predicted variable itself and the region in question (Haiden et al., 2024). These inaccuracies can propagate through hydrological models and lead to unreliable runoff forecasts, particularly under extreme conditions (Nester et al., 2012). To mitigate this issue, a variety of statistical and machine learning-based bias correction methods have been developed to adjust forecasted meteorological variables prior to their use as input in a hydrological model.</p>
      <p id="d2e148">A simple approach to reduce biases in precipitation is described by Lenderink et al. (2007), who scale precipitation linearly based on a constant factor calculated from long term observations. To support operational warning systems, Hess (2020) developed the Ensemble Model Output Statistics (Ensemble-MOS) system, which postprocesses ensemble forecasts from COSMO-D2-EPS and ECMWF-ENS. The approach relies on logistic regression and stepwise multiple regression to reduce conditional biases and produce calibrated probabilistic forecasts efficiently. Ko et al. (2020) used the XGBoost machine learning algorithm to correct precipitation forecasts. Their method demonstrates that machine learning can improve rainfall forecasting performance, especially localized heavy rainfall events, which are of special importance for flash floods in small catchments. Zhang et al. (2020) used LSTMs to learn relationships between meteorological forecasts and observed rainfall data. Their results indicate that LSTMs are capable of learning dynamic biases to correct the forecasts from numerical weather predictions and increase forecast reliability, especially for heavy rainfall events. Han et al. (2021) proposed CU-net, a convolutional neural network architecture specifically designed to address systematic biases in gridded numerical weather predictions from ECMWF-IFS. Their grid-based approach represents a methodological advancement by directly correcting spatial forecast fields, enabling comprehensive bias mitigation across continuous meteorological domains. However, the focus on ECMWF-IFS data raises important questions about the correction model's transferability to other numerical weather prediction systems, potentially limiting the generalizability of their bias correction framework to broader operational contexts.</p>
      <p id="d2e151">The studies mentioned have in common that the meteorological forecasts are compared either with meteorological station- or reanalysis data. In this context, it is important to note that especially precipitation measurements, whether from rain gauges, radar, or satellite sources, are inherently subject to various sources of uncertainty (Bárdossy et al., 2022). These errors stem from undercatch due to wind effects or sensor limitations (Yang et al., 1999). As a consequence, it can be assumed that even when inputting bias corrected precipitation forecast data to a hydrological model, a source of uncertainty with potential error propagation also arises here, which in turn creates a bias in runoff prediction. In contrast, discharge observations are typically regarded as more reliable compared to precipitation observations, as they represent an integrated hydrological response over the entire catchment and are measured continuously at fixed gauging stations (Herrnegger et al., 2015; Mao et al., 2019). Although discharge measurements also carry uncertainty, particularly related to the use of rating curves or sensor malfunction during extreme events, they are less affected by spatial representativeness errors, e.g. compared to precipitation, which requires spatial interpolation from a network of point measurements (De Oliveira and Vrugt, 2022; Villarini et al., 2008).</p>
      <p id="d2e155">A method directly improving streamflow forecasts from the physically based Global Flood Awareness System (GloFAS) was developed by Hunt et al. (2022). GloFAS is an operational hydrological forecasting system that couples ECMWF ensemble weather predictions with the LISFLOOD hydrological model to provide streamflow forecasts for rivers worldwide (Alfieri et al., 2013). Instead of bias-correcting the meteorological input variables, Hunt et al. (2022) addressed systematic biases in streamflow forecasts using a statistical bias correction method based on quantile mapping (QM) with spatial optimisation and subsequently applied a damping factor to blend the corrected forecasts with the original raw output. Despite the demonstrated improvements in forecast skill, this bias correction approach has several limitations. First, while the quantile mapping correction is not strictly limited to GloFAS and could in principle be applied to other distributed hydrological forecasting systems, it requires a hydrological forecast at the specific location of interest, which may not be available for all catchments, particularly smaller or poorly monitored ones. Second, the method is lead-time independent, meaning it does not account for the evolution of forecast bias over longer lead times, which can reduce its effectiveness for medium- to long-range forecasts. Third, the applied damping factor, while effective in reducing over-correction, is empirically tuned, which may limit its robustness when applied across diverse catchments or under changing climate conditions. A further limitation of the study is the relatively small number of catchments used (10 gauges), which constrains the generalizability of the findings.</p>
      <p id="d2e158">Building on the idea that runoff observations may be more accurate than those of meteorology, Kirchner (2009) proposed a paradigm shift through the concept of “doing hydrology backward”, where discharge is used as the primary constraint to infer the dynamics and uncertainties of upstream processes, such as precipitation or evapotranspiration. Rather than relying solely on uncertain meteorological inputs to predict runoff, backward hydrology extracts information about catchment dynamics directly from the discharge time series itself (Herrnegger et al., 2015; Kirchner, 2009). In this respect, the approach could also be used to perform a dynamic bias correction of multiple meteorological forecast variables since runoff data may serve as a more robust target variable in data-driven modelling frameworks than uncertain meteorological observations (e.g. rainfall). Given the hypothesis that large-scale hydrological datasets contain more information than could be described using theoretical or conceptual approaches (Nearing et al., 2021), a way to harness the potential of machine learning is to combine large sample datasets with meteorological forecasts as inputs. In such a setup, the model can learn to assign weights to the forecasts and internally correct their biases, thereby improving the overall runoff prediction accuracy.</p>
      <p id="d2e161">In this study we investigate multiple Long Short-Term Memory (LSTM) network architectures and training strategies to reduce meteorological forecast-induced bias in 24 h ahead maximum daily discharge predictions. The focus on daily maxima ensures that critical peak flows relevant to flood forecasting are not masked by temporal averaging. 24 h lead time was selected as an initial proof-of-concept to establish baseline performance of bias correction capabilities, as forecast uncertainty generally increases with lead time (Nester et al., 2012), making shorter horizons an appropriate starting point for validating the approach while providing a foundation for future extension to multi-day predictions. We evaluate baseline LSTM configurations, transfer learning approaches, encoder–decoder architectures, and sequential LSTM networks across 451 catchments from the Extended LamaH-CE dataset in Central Europe. Our experiments examine the effectiveness of different data integration scenarios, including the incorporation of past discharge observations and archived forecasts, with the goal of developing robust neural network-based approaches for operational flood forecasting systems that can effectively compensate for systematic biases inherent in numerical weather prediction models.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Data and Methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Data</title>
      <p id="d2e179">This study uses an extended version of the daily LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe (LamaH-CE; Klingler et al., 2021). LamaH consists of 859 gauged catchments including 21 catchment averaged meteorological variables, with more than 60 static catchment attributes. Since the original version of LamaH only contains meteorological ERA5-Land data and Kratzert et al. (2021) show that leveraging multiple meteorological data sources is beneficial in large sample hydrology, we expanded the data by 15 further variables from five sources. The products used are (i) ERA5-Land (Muñoz-Sabater et al., 2021) as in the original LamaH data, (ii) ECMWF-HRES European Center for Medium Range Weather Forecast - High Resoultion Forecast (ECMWF, 2025), (iii) E-OBS gridded observational data (Cornes et al., 2018), (iv) MSWEP multi-source weighted ensemble precipitation (Beck et al., 2019) and (v) GLEAM global land evaporation Amsterdam model (Miralles et al., 2011).  Details of the variables used, including their definitions, units, and sources, are summarized in Appendix A. The data products were obtained as raster data and subsequently aggregated to the LamaH basins. All variables are daily averages (e.g. temperature) or daily sums (e.g. precipitation). For the ECMWF-HRES variables temperature, dew point and sea level pressure, 3 hourly forecast values (8 per day) were calculated as daily averages starting from 0 o'clock (UTC) issue time. A second adaption we made to the LamaH dataset concerns the gauge files. In the daily version of LamaH-CE, there are only the mean daily discharges – we have extracted the daily minima and maxima from the hourly LamaH data for all gauges and extended the daily version with those.</p>

      <fig id="F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e184">LamaH domain with the 451 subset basins. For better illustration, LamaH subbasins (level B, blue polygons) are shown here, but calculations were performed at LamaH level A (lumped for each gauge).  The red points show the runoff gauges located at the catchment outlet.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f01.png"/>

        </fig>

      <p id="d2e193">For the conducted experiments, we used a subset of 451 basins with no and low anthropogenic influence at LamaH aggregation level A, which represents the lumped topographic catchment area of a gauge. Level A is comparable to the aggregation of the catchment areas in the CAMELS (Newman et al., 2015) dataset. The catchments are spatially distributed across the entire LamaH domain, with catchment areas including high alpine-, alpine foothill- and lowland areas. The subset includes both headwater and nested catchments, with 72.5 % of the basins being headwater catchments and 27.5 % being nested, meaning they have at least one other gauged basin from the selection upstream.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Experimental design</title>
      <p id="d2e204">To comprehensively evaluate the performance of LSTM-based flood prediction models under different data availability scenarios and training strategies, we designed five distinct experimental groups with the primary research question: <italic>How to reduce the meteorological forecast induced bias in runoff predictions?</italic></p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e211">Workflow of the conducted experiments. The LamaH-CE dataset was extended by forecast and further reanalysis data and subset to 451 basins with no and low anthropogenic influence. The experimental split is divided into the deep learning architectures used, followed by a schematic representation of input embeddings for static and dynamic variables which feature space is fed to the LSTM. The last step is the forecast of the daily maximum runoff at the gauges.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f02.png"/>

        </fig>

      <p id="d2e220">All experiments were, in terms of reproducibility, conducted with the NeuralHydrology (Kratzert et al., 2022) python library and trained on different LSTM architectures to predict maximum daily discharge (<inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>). The models incorporated dynamic meteorological inputs using a 365 d input sequence length, and static catchment attributes (33 physiographic, climatic, and land cover characteristics), both processed through separate embedding networks. The embedding networks are fully connected neural network layers that transform raw input variables into learned representations (Ahmed et al., 2023). A primary motivation for their use in this study is to enable transfer learning and consistent processing across experiments with different input dimensions. By projecting varying numbers of input variables from different data sources into a shared latent space, the model can be pre-trained on reanalysis and fine-tuned on forecast data regardless of differences in the feature space between domains. The experimental framework utilized a consistent temporal split with training data from 2003–2009, validation from 2010–2013, and testing from 2014–2017. Model performance was evaluated using basin averaged Nash-Sutcliffe Efficiency  (NSE*, Kratzert et al., 2019c) as the primary loss function. A description of the loss function is attached in Appendix C. Model hyperparameters, such as the number of hidden units, were optimized using Bayesian optimization (see Snoek et al., 2012) with NSE* as the objective function. A detailed description of the performed hyperparameter tuning is attached in Appendix D. Hereafter, we use the term “domain” in its machine learning sense, referring to a specific data distribution characterized by its feature space and statistical properties, rather than in its common hydrological meaning of a geographical region. Throughout the remainder of this paper, domain specifications and equations reference only dynamic meteorological forcings for clarity, with the understanding that static catchment attributes remain unchanged across all experimental setups.</p>
<sec id="Ch1.S2.SS2.SSS1">
  <label>2.2.1</label><title>Baseline</title>
      <p id="d2e242">Three baseline experiments were conducted to establish performance benchmarks using different meteorological data sources in a standard LSTM runoff simulation framework. LSTMs are a special form of recurrent neural networks, mainly used for sequential (time series) data (Hochreiter and Schmidhuber, 1997). For a detailed description of the LSTM in relation to hydrological modelling, we refer to Kratzert et al. (2018, 2019b, c). The core tensor equations of the LSTM model responsible for the information flow are presented in Appendix E.</p>
      <p id="d2e245">The baseline experiments were conducted using either forecasting data only (FC), reanalysis data only (RA), or a combination of both (FCRA) as dynamic inputs. Following Seibert et al. (2018), who argue that model performance can only be meaningfully interpreted when evaluated relative to benchmarks representing what could and should be expected, the FC experiment, forced only with archived forecasting data from ECMWF HRES, serves as a lower benchmark. Here, the 1 d lead-time ECMWF-HRES forecast is used for the full 365 d input sequence. The RA experiment, exclusively driven by all available reanalysis and observational data sources (ERA5-Land, E-OBS, MSWEP, and GLEAM), as described in Appendix A, establishes an upper benchmark for model performance under ideal hindcast conditions (i.e. retrospective simulations using quality-controlled historical data). The FCRA experiment extended the RA experiment by additionally incorporating the five ECMWF-HRES forecast variables as dynamic inputs alongside all reanalysis and observational data sources, representing the optimal data availability scenario.</p>
      <p id="d2e248">The domains in the experiments formulate as:

                  <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M2" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E1"><mml:mtd><mml:mtext>1</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E2"><mml:mtd><mml:mtext>2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mfenced open="(" close=")"><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E3"><mml:mtd><mml:mtext>3</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FCRA</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            <inline-formula><mml:math id="M3" display="inline"><mml:mi mathvariant="script">D</mml:mi></mml:math></inline-formula>: Dataset used in the experiment; <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>: Meteorological variables at timestep <inline-formula><mml:math id="M5" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>; <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>: Maximum daily discharge (target variable) at timestep <inline-formula><mml:math id="M7" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS2">
  <label>2.2.2</label><title>Cross-Domain Evaluation</title>
      <p id="d2e470">The cross domain evaluation (CD) experiment examined model generalization by training on reanalysis data and testing on ECMWF-HRES forecast data while maintaining 5 identical input variables. The meteorological variables used in this experiment are temperature, dewpoint temperature, precipitation, solar radiation and actual evapotranspiration. This experimental design mirrors the operational framework of classic conceptual hydrological models, where models are typically calibrated using high-quality reanalysis data with subsequently applied real-time forecast inputs during operational usage. By replicating this established modelling paradigm within the LSTM framework, the experiment quantifies the performance shift when transitioning from reanalysis to operationally available forecast data. It is important to note that the CrossDomain experiment deliberately uses only five reanalysis variables, those with a direct equivalent in the ECMWF-HRES forecast data, in contrast to the Baseline Reanalysis which uses all 31 available reanalysis variables. The performance gap between these two configurations in Fig. 3 therefore reflects the additional benefit of combining multiple meteorological data sources, consistent with Kratzert et al. (2021). Our hypothesis here was that if the distributions of the two input data sets <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">RA</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">CrossDomain</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> are too different, the model performance will decline.</p>
      <p id="d2e500">The variables compared in the Cross Domain Evaluation can be taken from Table 1, with its domains in the experiments formulated as:

                  <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M10" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E4"><mml:mtd><mml:mtext>4</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>Train on: </mml:mtext><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">RA</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">CrossDomain</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E5"><mml:mtd><mml:mtext>5</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtext>Test on: </mml:mtext><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mfenced open="(" close=")"><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:msubsup><mml:mo mathvariant="italic">}</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula></p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e618">The five meteorological forcing variables shared between reanalysis and forecast, used in the Cross Domain Evaluation experiment to ensure an identical input feature space between training and inference. Variables were deliberately sourced from multiple dedicated products to combine the strengths of different data sources, reduce cross-correlation between input features, and avoid propagating systematic biases from a single reanalysis across all variables. Especially for total precipitation, MSWEP has shown to consistently outperform other precipitation datasets across diverse hydroclimatic settings (Abbas et al., 2026; Beck et al., 2017). A detailed description of all variables is provided in Appendix A.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:colspec colnum="6" colname="col6" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Forcing</oasis:entry>
         <oasis:entry colname="col2">Temperature</oasis:entry>
         <oasis:entry colname="col3">Precipitation</oasis:entry>
         <oasis:entry colname="col4">Solar Radiation</oasis:entry>
         <oasis:entry colname="col5">Dewpoint Temp</oasis:entry>
         <oasis:entry colname="col6">Actual</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Source</oasis:entry>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4"/>
         <oasis:entry colname="col5"/>
         <oasis:entry colname="col6">Evapotranspiration</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Forecast</oasis:entry>
         <oasis:entry colname="col2">ECMWF_t2m</oasis:entry>
         <oasis:entry colname="col3">ECMWF_tp</oasis:entry>
         <oasis:entry colname="col4">ECMWF_ssrd</oasis:entry>
         <oasis:entry colname="col5">ECMWF_d2m</oasis:entry>
         <oasis:entry colname="col6">ECMWF_e</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Reanalysis</oasis:entry>
         <oasis:entry colname="col2">ERA5L_2m_temp_mean</oasis:entry>
         <oasis:entry colname="col3">MSWEP_RR</oasis:entry>
         <oasis:entry colname="col4">EOBS_qq</oasis:entry>
         <oasis:entry colname="col5">ERA5L_2m_dp_temp_mean</oasis:entry>
         <oasis:entry colname="col6">GLEAM_ETA</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S2.SS2.SSS3">
  <label>2.2.3</label><title>Encoder–Decoder LSTM</title>
      <p id="d2e737">The Encoder–Decoder LSTM developed by Nearing et al. (2024) consists of two connected LSTMs: one for the hindcast phase forced with historical meteorological reanalysis data (e.g. ERA5) and one for the forecast phase forced with weather forecast data. The two LSTMs are connected by a non-linear handoff network in which the cell state and hidden state from the hindcast are transferred to the forecast LSTM. This architectural design allows the forecast LSTM to learn hydrological states from the hindcast, which could be understood as initial conditions in the model.</p>
      <p id="d2e740">Three distinct experiments were implemented using the Encoder–Decoder LSTM architecture to investigate if this dual-LSTM framework can learn and compensate dynamical biases inherent in meteorological forecasts. The first experiment with domain <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> implemented the basic encoder–decoder framework where the hindcast LSTM was forced with historical reanalysis data while the forecast LSTM processed meteorological forecast data. The second experiment with domain <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> extended this architecture by incorporating past mean daily discharge observations alongside reanalysis data in the hindcast LSTM. The third experiment with domain <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> extended <inline-formula><mml:math id="M14" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> by additionally forcing the hindcast cell of the model with forecast data. This emulates a setting in which archived forecasting data are used in combination with reanalysis data in the hindcast phase.</p>
      <p id="d2e823">The domains in the experiments formulate as:

                  <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M15" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E6"><mml:mtd><mml:mtext>6</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mfenced close="}" open="{"><mml:mfenced open="(" close=")"><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mfenced><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E7"><mml:mtd><mml:mtext>7</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:msubsup><mml:mfenced close="}" open="{"><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>∪</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mfenced><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E8"><mml:mtd><mml:mtext>8</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>=</mml:mo><mml:msubsup><mml:mfenced open="{" close="}"><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>∪</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mfenced><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula>

            <inline-formula><mml:math id="M16" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula>: Sequence length.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS4">
  <label>2.2.4</label><title>Sequential Forecast LSTM</title>
      <p id="d2e1182">The Sequential Forecast LSTM experiment employs a two-phase sequential processing strategy to leverage both reanalysis and operationally available forecast data within a unified framework. The architecture consists of separate embedding networks for hindcast and forecast inputs, a shared LSTM layer and a state transfer mechanism that enables knowledge transfer between processing phases (see Sequential Forecast LSTM in NeuralHydrology, Kratzert et al., 2022). In the first phase, the LSTM processes embedded historical reanalysis data to generate hidden and cell states. The second phase continues LSTM processing with embedded forecast data, initialized with the states from the hindcast phase, ensuring that forecast predictions are informed by contextual information learned from historical patterns. The model generates predictions by concatenating outputs from both phases through a prediction head, with the optimization objective to maximize NSE*. This design enables optimal utilization of reanalysis data for learning hydrological patterns while maintaining operational forecasting capabilities through the state transfer mechanism.</p>
      <p id="d2e1185">Experiment one (<inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) used the basic Sequential LSTM framework, with only using reanalysis data in the hindcast phase and forecast data in the forecast phase. The second experiment (<inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>) added to the first domain mean daily discharge observations alongside reanalysis data in the hindcast phase. In the third experiment (<inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>), we extended <inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> by additionally forcing the hindcast phase of the model with archived forecast data.</p>
      <p id="d2e1252">The domains in the experiments formulate as:

                  <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M21" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E9"><mml:mtd><mml:mtext>9</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mfenced close="}" open="{"><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mfenced><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E10"><mml:mtd><mml:mtext>10</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace width="0.25em" linebreak="nobreak"/><mml:mo>=</mml:mo><mml:msubsup><mml:mfenced close="}" open="{"><mml:mfenced close=")" open="("><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>∪</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mfenced><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E11"><mml:mtd><mml:mtext>11</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtable class="split" rowspacing="0.2ex" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mo>=</mml:mo><mml:msubsup><mml:mfenced close="}" open="{"><mml:mfenced open="(" close=")"><mml:mrow><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi mathvariant="normal">RA</mml:mi></mml:msubsup><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:mi>F</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msubsup><mml:mo>∪</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>s</mml:mi><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>∪</mml:mo><mml:msubsup><mml:mi>x</mml:mi><mml:mi>t</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msubsup><mml:mo>,</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi mathvariant="normal">max</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mfenced><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mi>s</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="Ch1.S2.SS2.SSS5">
  <label>2.2.5</label><title>Transfer Learning</title>
      <p id="d2e1594">Transfer Learning (TL) is a machine learning paradigm leveraging gained knowledge from a source domain to improve learning performance in a target domain (Goodfellow et al., 2016). Formally, TL aims to improve the predictive performance on the target domain using knowledge from the source domain, with differences potentially existing in the feature space, data distribution, or learning task between the two domains (Zhuang et al., 2021). TL can be categorized into two primary types based on the relationship between source and target domain: The first is homogeneous transfer learning, where both domains share the same feature space (i.e. using identical meteorological variables and catchment attributes) and have the same marginal probability distributions (Weiss et al., 2016). The second is heterogeneous transfer learning, where the feature spaces differ between domains (Pan and Yang, 2010). For our experiments, we used the heterogeneous transfer learning approach – while the learning task stays the same in the conducted experiments, namely predicting maximum daily discharges at a gauge, the feature spaces and its distributions between forecast (target domain) and reanalysis (source domain) data differs, as evidenced by the violin plots in Appendix B.</p>
      <p id="d2e1597">In the context of forecast bias reduction, transfer learning is used to leverage knowledge from the less bias-influenced reanalysis source data to improve prediction accuracy when applied to the more bias-prone forecast target data. This approach is particularly relevant in contexts involving hydrometeorological data, where reanalysis data represents a post-processed quality-controlled dataset with reduced systematic errors, while forecast data contains dynamical biases from numerical weather prediction models. By pre-training the temporal encoder (LSTM) on reanalysis data, the model learns hydrological process representations that can subsequently be fine-tuned to accommodate the bias characteristics of forecast inputs, potentially improving the model's ability to correct for systematic forecast errors while maintaining learned temporal dependencies.</p>
      <p id="d2e1600">The first experiment implemented full weight transfer learning, where all network weights from the baseline <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> experiment (embedding networks, LSTM and output layers pre-trained on reanalysis data) were used as initialization for a new training phase on forecast data, allowing all parameters to be updated through backpropagation to adapt to the <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> target domain's characteristics. The second experiment, also based on the weights of  <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> employed selective weight transfer learning, adapting only the dynamic embedding network weights while freezing other model parameters including the static embedding network, thus preserving learned temporal patterns while allowing adaptation to new forcing input characteristics. The third experiment applied the same selective transfer learning method as the second experiment, but network weights are based on the <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FCRA</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> domain.</p>
      <p id="d2e1647">The domains in the experiments are given below with source and target domains as well as training objective. All three experiments are based on training either LSTM weights <inline-formula><mml:math id="M26" display="inline"><mml:mi mathvariant="italic">θ</mml:mi></mml:math></inline-formula>, embedding layer weights <inline-formula><mml:math id="M27" display="inline"><mml:mi mathvariant="italic">ϕ</mml:mi></mml:math></inline-formula>, output layer weights <inline-formula><mml:math id="M28" display="inline"><mml:mi mathvariant="italic">ψ</mml:mi></mml:math></inline-formula> or all combined with a NSE* loss function <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">NSE</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>.

                  <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M30" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="Ch1.E12"><mml:mtd><mml:mtext>12</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mtext>TL</mml:mtext><mml:mi mathvariant="normal">AllWeights</mml:mi></mml:msub><mml:mo>:</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext> source domain </mml:mtext><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mtext> target domain </mml:mtext><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:mtext> with objective </mml:mtext><mml:mi mathvariant="normal">arg</mml:mi><mml:munder><mml:mi mathvariant="normal">min</mml:mi><mml:mrow><mml:mi mathvariant="italic">θ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ϕ</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ψ</mml:mi></mml:mrow></mml:munder><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">NSE</mml:mi></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E13"><mml:mtd><mml:mtext>13</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mtext>TL</mml:mtext><mml:mrow><mml:mi mathvariant="normal">EmbeddingNet</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext> source domain </mml:mtext><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mtext> target domain </mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub><mml:mtext> with objective </mml:mtext><mml:mi mathvariant="normal">arg</mml:mi><mml:munder><mml:mi mathvariant="normal">min</mml:mi><mml:mo mathvariant="italic">ϕ</mml:mo></mml:munder><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">NSE</mml:mi></mml:msub><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="Ch1.E14"><mml:mtd><mml:mtext>14</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:msub><mml:mtext>TL</mml:mtext><mml:mrow><mml:mi mathvariant="normal">EmbeddingNet</mml:mi><mml:mi mathvariant="normal">_</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub><mml:mo>:</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mtext> source domain </mml:mtext><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FCRA</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:mtext> target domain </mml:mtext></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub><mml:mtext> with objective </mml:mtext><mml:mi mathvariant="normal">arg</mml:mi><mml:munder><mml:mi mathvariant="normal">min</mml:mi><mml:mo mathvariant="italic">ϕ</mml:mo></mml:munder><mml:msub><mml:mi>L</mml:mi><mml:mi mathvariant="normal">NSE</mml:mi></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula></p>
</sec>
<sec id="Ch1.S2.SS2.SSS6">
  <label>2.2.6</label><title>Input Embedding</title>
      <p id="d2e1907">Input embedding networks serve as pre-processing layers that transform raw meteorological variables into fixed dimensional representations for LSTM processing. The embedding layers enable the model to learn (non-) linear combinations and scaling of input features, potentially capturing complex relationships between meteorological variables that may not be apparent in their original form (Irani et al., 2025). The embedding transformation is relevant for hydrometeorological applications where variables such as temperature, precipitation and solar radiation may exhibit non-linear interactions that influence runoff generation processes. To investigate the impact of embedding complexity on bias correction performance, we implemented two distinct embedding architectures: a simple embedding consisting of a single fully connected layer with 16 hidden units and tanh activation, and a complex embedding featuring a three-layer network with 30, 20, and 64 hidden units respectively, also using tanh activation functions. The simple embedding provides a lightweight transformation with minimal parameter overhead, while the complex embedding offers greater representational capacity through deeper non-linear transformations.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Quantifying Domain Shift and Its Physiographic Controls</title>
      <p id="d2e1919">The premise of this study is that the distributional differences between reanalysis and forecast meteorological data, as visually evident in the violin plots of Appendix B, lead to substantial performance degradation when transitioning between the two domains. To investigate the spatial patterns of this domain shift and their relationship to catchment characteristics, two complementary analyses are conducted. First, the 1-Wasserstein distance (Villani, 2009) is computed for each of the five shared meteorological variables between the reanalysis and forecast distributions across all 451 basins during the test period (2014–2017), using the input domains from the cross-domain evaluation experiment. The 1-Wasserstein distance measures the minimum effort required to transform one probability distribution into another and is computed in the original physical units of each variable to preserve interpretability. It is defined as:

            <disp-formula id="Ch1.E15" content-type="numbered"><label>15</label><mml:math id="M31" display="block"><mml:mrow><mml:msub><mml:mi>W</mml:mi><mml:mn mathvariant="normal">1</mml:mn></mml:msub><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>P</mml:mi><mml:mo>,</mml:mo><mml:mi>Q</mml:mi></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:munderover><mml:mo movablelimits="false">∫</mml:mo><mml:mrow><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow><mml:mrow><mml:mo>+</mml:mo><mml:mi mathvariant="normal">∞</mml:mi></mml:mrow></mml:munderover><mml:mfenced close="|" open="|"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:msub><mml:mfenced open="(" close=")"><mml:mi>x</mml:mi></mml:mfenced><mml:mo>-</mml:mo><mml:msub><mml:mi>F</mml:mi><mml:mi>Q</mml:mi></mml:msub><mml:mfenced open="(" close=")"><mml:mi>x</mml:mi></mml:mfenced></mml:mrow></mml:mfenced><mml:mspace width="0.125em" linebreak="nobreak"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>P</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi>F</mml:mi><mml:mi>Q</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are the empirical cumulative distribution functions of the reanalysis and forecast distributions respectively. The resulting per-basin Wasserstein distances are then correlated with the 33 static catchment attributes listed in Table A2 using Spearman correlation and visualized as a heatmap to identify which physiographic characteristics are associated with larger meteorological domain shifts.</p>
      <p id="d2e2003">Further, to investigate how catchment characteristics modulate model skill across all experimental configurations a two- fold analysis is conducted. First, all basins of the experimental domains <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> are compared against those of the Forecast Baseline <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi>F</mml:mi><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>by calculating the per basins difference of the NSE values to assess in which basins the model performance possibly increases or decreases – defined as:

            <disp-formula id="Ch1.E16" content-type="numbered"><label>16</label><mml:math id="M39" display="block"><mml:mtable rowspacing="0.2ex" class="split" displaystyle="true" columnalign="right left"><mml:mtr><mml:mtd><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="normal">NSE</mml:mi><mml:mo>=</mml:mo></mml:mrow></mml:mtd><mml:mtd><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">ED</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">LSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:mtd></mml:mtr><mml:mtr><mml:mtd/><mml:mtd><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mrow><mml:mi mathvariant="normal">SEQLSTM</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>

          Inspired by the work of Seibert et al. (2018), suggesting to compare hydrological model experiments against well-defined benchmarks, we introduce a data plot technique to compare experimental differences in large sample hydrology settings: Plotting a heatmap with the model experiments as rows, <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="normal">NSE</mml:mi></mml:mrow></mml:math></inline-formula> per basin as columns and sorting the columns by the mean <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="normal">NSE</mml:mi></mml:mrow></mml:math></inline-formula> over all experiments per basin (columns). This plot allows a simple visual comparison of all experimental differences per basin and enables detecting outliers. While Seibert et al. (2018) propose a simple conceptual bucket-type model as the lower benchmark, we use the Baseline Forecast as the lower reference because it represents the most operationally relevant starting point for our specific problem: a LSTM model trained and tested exclusively on operational ECMWF HRES forecast data, without any bias correction strategy applied.</p>
      <p id="d2e2197">Second, the Spearman correlation between all positive per-basin <inline-formula><mml:math id="M42" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE and all 33 static catchment attributes is computed for each experiment and visualized as a heatmap. The focus on positive error metric differences (<inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi mathvariant="normal">NSE</mml:mi><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>) was chosen since it indicates those basins where the proposed strategy outperforms the Baseline Forecast, effectively capturing the physiographic controls on forecast skill improvement. It is important to note that higher positive <inline-formula><mml:math id="M44" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE values do not exclusively reflect bias reduction – they also incorporate improvements stemming from increased model complexity, such as the additional architectural components and discharge integration introduced in the more advanced configurations.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Results and Discussion</title>
      <p id="d2e2237">All results presented in subsequent sections are visualized as Cumulative Density Functions (CDFs) of Nash-Sutcliffe Efficiency values computed across the 451 study basins, where each point represents the proportion of basins achieving a specified NSE value. This visualization approach enables comprehensive assessment of model performance distribution, revealing not only median performance but also the full range of model behaviour across diverse catchment conditions. The light grey line denotes the forecast-only baseline <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">FC</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula> representing the lower performance bound, while the dark grey line denotes the reanalysis-only baseline <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula> establishing the upper performance bound. These reference curves remain consistent across all figures to facilitate direct comparison between experimental configurations. A comprehensive overview of all performance statistics including median, mean, standard deviation and percentiles for each experimental configuration is provided in Appendix F.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Propagation of Forecast Uncertainty in the Hydrological Model Setting</title>
      <p id="d2e2274">A fast and straightforward way to analyse the propagation of forecast uncertainty in predicting maximum daily discharge (<inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) is the cross-domain evaluation (CD) experiment, depicted in Fig. 3. The CD experiment replicates the classical hydrological modelling workflow, where a model is calibrated on reanalysis data and subsequently applied with forecast inputs. To ensure that the input feature space remains identical between training and inference, only the five variables that have a direct equivalent in both reanalysis and forecast are used (ERA5L_2m_temp_mean, ERA5L_2m_dp_temp_mean, ERA5L_surf_net_solar_rad_mean, MSWEP_RR, GLEAM_ETA), in contrast to the Baseline Reanalysis model which uses all 31 available reanalysis variables from ERA5-Land, E-OBS, MSWEP, and GLEAM. CD reveals degradation in hydrological model performance when transitioning from reanalysis to forecast meteorological forcings, despite five identical input variables. The median Nash-Sutcliffe Efficiency (NSE) decreased from 0.58 to 0.33, representing a 0.25 reduction in model skill. This performance deterioration is accompanied by increased uncertainty, with the NSE standard deviation rising from 0.87 to 1.1, indicating that forecast uncertainty propagates through the hydrological model shifting and broadening the NSE distribution. The mean NSE exhibits an even more pronounced decline (0.44 to 0.19), suggesting increased negative skewness due to extreme poor-performing outliers.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e2290">Cumulative Density Function of Nash Sutcliffe Efficiency values for the Cross Domain Evaluation Experiment. The comparison includes the Pre trained model on five meteorological reanalysis variables (dark blue), the One Shot (direct application without fine-tuning) based on the weights of the Pre trained model applied to the corresponding five meteorological forecasting variables from ECMWF-HRES and the baselines (grey). The vertical dashed lines depict the median NSE for each experiment. The blue arrow shows the performance decrease at median NSE when applying cross domain evaluation.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f03.png"/>

        </fig>

      <p id="d2e2299">The underlying cause of this performance decrease can be attributed to the difference in data distributions between reanalysis and forecast datasets (see Appendix B), effectively representing a domain shift problem. Domain shift occurs when the statistical properties of the training data (reanalysis) differ from those of the target data (forecast), violating the fundamental assumption of independent and identically distributed data that underlies machine learning model generalization (Goodfellow et al., 2016; Hosna et al., 2022). Neural networks are particularly susceptible to domain shifts as they learn to map input-output relationships based on the specific distributional characteristics of their training data, leading to degraded performance when deployed on data from a different distribution. These results demonstrate that it is not feasible to simply substitute reanalysis data with forecast data in neural network-based hydrological modeling applications, as meteorological forecast uncertainty propagates through the model chain, degrading the representation of catchment processes and creating performance risks for operational hydrological forecasting systems.</p>
<sec id="Ch1.S3.SS1.SSSx1" specific-use="unnumbered">
  <title>Linking Meteorological Domain Shift to Topographical Features</title>
      <p id="d2e2309">Having established that the performance degradation in the cross-domain evaluation is driven by distributional differences between reanalysis and forecast inputs, we now investigate where across the 451 basins these differences are largest and which catchment characteristics are associated with larger domain shifts.</p>
      <p id="d2e2312">The Spearman correlation between the per-basin 1-Wasserstein distances and the 33 static catchment attributes reveals consistent patterns across all five meteorological variables, as shown in Fig. 4. Elevation-related attributes emerge as a strong and consistent predictor of domain shift across all variables, with mean elevation showing positive correlations ranging from <inline-formula><mml:math id="M48" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.19</mml:mn></mml:mrow></mml:math></inline-formula> (dewpoint temperature) to <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.67</mml:mn></mml:mrow></mml:math></inline-formula> (precipitation), indicating that high-elevation catchments experience consistently larger distributional differences between reanalysis and ECMWF-HRES. This pattern is physically plausible, as ECMWF-HRES forecast skill is generally lower in complex alpine terrain where orographic effects are difficult to resolve at the model resolution (Haiden et al., 2024; Lavers et al., 2021).</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e2341">Spearman correlation between the per-basin 1-Wasserstein distances of five shared meteorological variables (Surface Solar Radiation, 2 m Temperature, 2 m Dewpoint Temperature, Precipitation, Actual Evapotranspiration) and 33 static catchment attributes across all 451 basins during the test period (2014–2017). Positive values (red) indicate that catchments with larger values of a given attribute tend to experience larger distributional differences between ERA5-Land reanalysis and ECMWF-HRES forecast inputs, while negative values (blue) indicate smaller distributional differences. Wasserstein distances are computed in the original physical units of each variable using the domains from the cross-domain evaluation experiment.</p></caption>
            <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f04.png"/>

          </fig>

      <p id="d2e2350">For surface solar radiation, the strongest positive correlation is found with bare fraction (<inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.54</mml:mn></mml:mrow></mml:math></inline-formula>) and mean elevation (<inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.51</mml:mn></mml:mrow></mml:math></inline-formula>), while the strongest negative correlation is with mean actual evapotranspiration (<inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.63</mml:mn></mml:mrow></mml:math></inline-formula>) and LAI maximum (<inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.60</mml:mn></mml:mrow></mml:math></inline-formula>), indicating that vegetated, low-elevation catchments with high evapotranspiration experience the smallest domain shift for this variable. For precipitation, mean elevation (<inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.67</mml:mn></mml:mrow></mml:math></inline-formula>) and slope mean (<inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.63</mml:mn></mml:mrow></mml:math></inline-formula>) show the strongest positive correlations, while aridity index shows the strongest negative correlation (<inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.66</mml:mn></mml:mrow></mml:math></inline-formula>), suggesting that steeper, higher-elevation and wetter catchments experience larger distributional differences between the two precipitation products. For temperature, aridity index shows the strongest negative correlation (<inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.60</mml:mn></mml:mrow></mml:math></inline-formula>), indicating that wetter catchments experience larger distributional differences between the two products, consistent with the positive correlations found for mean precipitation (<inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.57</mml:mn></mml:mrow></mml:math></inline-formula>), fraction of snow (<inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.56</mml:mn></mml:mrow></mml:math></inline-formula>) and slope mean (<inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.52</mml:mn></mml:mrow></mml:math></inline-formula>), which are positively associated with wetter and more topographically complex conditions. For dewpoint temperature, correlations are substantially weaker across all attributes, with precipitation seasonality (<inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.31</mml:mn></mml:mrow></mml:math></inline-formula>) and vegetation-related attributes such as GVF maximum (<inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.38</mml:mn></mml:mrow></mml:math></inline-formula>) emerging as the dominant controls. For actual evapotranspiration, mean precipitation (<inline-formula><mml:math id="M63" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.58</mml:mn></mml:mrow></mml:math></inline-formula>) and fraction of snow (<inline-formula><mml:math id="M64" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.52</mml:mn></mml:mrow></mml:math></inline-formula>) show the strongest positive correlations, while aridity index (<inline-formula><mml:math id="M65" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.55</mml:mn></mml:mrow></mml:math></inline-formula>) and agricultural fraction (<inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.53</mml:mn></mml:mrow></mml:math></inline-formula>) show the strongest negative correlations, again suggesting that agriculturally dominated and drier catchments experience smaller domain shifts.</p>
      <p id="d2e2574">Overall, the results consistently indicate that topographically complex, high-elevation, snow-dominated catchments experience the largest distributional differences between reanalysis and forecast inputs, while low-elevation and vegetated catchments show smaller domain shifts.</p>
</sec>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Performance Analysis Across Unmodified Architectures</title>
      <p id="d2e2586">To address the domain shift challenges identified in Sect. 3.1, we evaluated different neural network architectures and training techniques, presenting here the optimal configurations from each experimental setup.</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e2591">Cumulative Density Function of Nash Sutcliffe Efficiency values for the Unmodified Architectures Experiment. The experiments include the baselines (grey), comprehensive transfer learning with a finetuning of the embedding and LSTM weights (TL AllWeights), selective transfer learning with only finetuning the embedding weights while the LSTM was frozen (TL EmbeddingNet), Encoder- Decoder LSTM (purple) and the Sequential Forecast LSTM (orange). The vertical dashed lines depict the median NSE for each experiment. The arrows indicate the change in median NSE values relative to the baseline forecast experiment.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f05.png"/>

        </fig>

      <p id="d2e2600">Transfer learning was implemented through two contrasting approaches: comprehensive parameter updating, where the entire network was retrained on forecast data (TL AllWeights), and selective embedding retraining, where only the input embedding layers were fine-tuned while maintaining the pre-trained weights of the deeper network components (TL EmbeddingNet). Both experiments utilized the reanalysis baseline with a complex embedding layer as the starting point. The selective embedding approach (TL EmbeddingNet) achieved higher performance, with a median NSE of 0.44, representing a 0.11 improvement over the cross-domain baseline (0.33). The advantage of selective retraining becomes evident when comparing the two transfer learning strategies: TL EmbeddingNet shows an improvement in the 10th percentile (0.15) compared to TL AllWeights (0.05), indicating that the retraining approach is especially effective for poorly performing basins. The enhanced performance stems from the model's ability to leverage robust hydrometeorological relationships learned from reanalysis data while adapting the input representation to forecast data characteristics. Reanalysis products provide more complete and physically consistent atmospheric descriptions through data assimilation, enabling the model to learn generalized process representations that are subsequently refined during fine-tuning on forecast data.</p>
      <p id="d2e2604">The Encoder–Decoder LSTM architecture demonstrated performance improvements over the transfer learning approaches, achieving a median NSE of 0.57. This architecture showed particular strength in the upper performance range, with the 75th and 90th percentiles reaching 0.65 and 0.73, respectively.</p>
      <p id="d2e2607">The Sequential Forecast LSTM achieved the highest overall performance among the unmodified forecast-based configurations, with a median NSE of 0.63 and notably consistent results across all percentiles. The model demonstrated higher stability compared to other approaches, evidenced by the low standard deviation (0.52). Among the unmodified architectures, the Sequential LSTM is best able to efficiently reduce forecasting bias. We attribute this to its two-phase processing structure, where the hindcast and forecast phases are processed consecutively within the same LSTM, maintaining a single continuous state evolution from hindcast to forecast. The cell state serves as long-term memory that can selectively retain or forget information across time steps, while the hidden state captures current relevant information, enabling the model to maintain both short-term adaptations to recent forecast patterns and long-term memory of systematic biases. Unlike the Encoder- Decoder architecture transforming hindcast states through a fixed handoff network before initiating forecast processing, the Sequential LSTM maintains continuous state evolution from hindcast to forecast, preserving temporal patterns without disruption and potentially enabling better compensation for systematic errors that emerge at the forecast transition.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>The role of integrating archived forecasts in the hindcast phase</title>
      <p id="d2e2619">Given the performance degradation observed when transitioning from reanalysis to forecast data, we investigated whether training models with a combination of reanalysis and archived forecast data could improve forecast performance. In short, the integration of archived forecasts in the hindcast phase demonstrates limited effectiveness across all tested architectures.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e2624">Cumulative Density Function of Nash Sutcliffe Efficiency values for experiments integrating archived forecasts in the hindcast phase. The experiments include the baselines (grey), selective transfer learning with only finetuning the embedding weights while the LSTM was freezed (TL EmbeddingNet), Encoder- Decoder LSTM (purple) and the Sequential Forecast LSTM (orange). Solid lines indicate the experiments with solely reanalysis data in the hindcast phase, while dashed lines display the combination of archived forecasts and reanalysis in the hindcast phase.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f06.png"/>

        </fig>

      <p id="d2e2633">In Fig. 5 it is evident that in transfer learning only fine-tuning the embedding net led to higher forecast skill, leading us to only focus in this experimental setup on the selective transfer learning method. As previously discussed, the TL EmbeddingNet approach pretrained exclusively on reanalysis data achieved a median NSE of 0.44. In contrast, when the same transfer learning architecture is pretrained on combined forecast and reanalysis data, the results show deteriorated performance compared to the reanalysis-only pretraining, with a median NSE dropping to 0.41. This configuration exhibits increased variability (NSE standard deviation of 2.29) while the 10th percentile performance remains poor at 0.02, failing to achieve the improvement (0.15) observed with reanalysis-only pretraining. The mean performance also deteriorates significantly (0.21 vs. 0.35 for reanalysis-only), indicating the introduction of more negative outliers.</p>
      <p id="d2e2637">For the Encoder–Decoder LSTM, incorporating archived forecasts yields marginal improvements in the upper performance percentiles, with the 75th and 90th percentiles increasing from 0.65 to 0.69 and 0.73 to 0.77, respectively, while the 10th percentile improves slightly from 0.21 to 0.24. However, these modest gains are accompanied by increased variability (standard deviation increases from 8.03 to 9.33) while the median NSE remains unchanged at 0.57.</p>
      <p id="d2e2640">The Sequential Forecast LSTM shows no meaningful benefit from archived forecast integration, with the median NSE decreasing marginally from 0.63 to 0.62. More critically, this architecture experiences an increase in variability (standard deviation from 0.52 to 6.28) and the mean performance drops from 0.57 to 0.19, indicating the introduction of numerous extreme negative outliers that significantly compromise model reliability. A systematic explanation for these outliers has not been identified, but individual inspection of specific basins suggests that residual anthropogenic influences may play a role. For example, Basin 758 shows a persistent and strong overestimation of simulated discharge relative to observations throughout the test period, which is consistent with the presence of water abstractions, retention structures, or other human interventions not represented in the model. Despite filtering the 451 basins for low anthropogenic influence based on the LamaH-CE catchment classification, some residual human influence cannot be excluded.</p>
      <p id="d2e2643">The results demonstrate that integrating archived forecasts during the hindcast phase does not provide meaningful performance improvements for either architecture. The approach either yields negligible benefits while increasing instability (Encoder–Decoder LSTM) or actively degrades performance (Transfer learning, Sequential LSTM), indicating that this strategy is not effective for addressing domain shift challenges in hydrological forecasting applications.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>The role of integrating past discharge in the training domain</title>
      <p id="d2e2654">Previous studies have consistently focused on modeling ungauged basins (Kratzert et al., 2019b; Nearing et al., 2024). However, we argue that when near real-time discharge data are available, as is the case for most gauging stations across the LamaH catchments in Central Europe, the incorporation of discharge observations positively influences forecast accuracy. For example, in the Austrian LamaH basins, discharge data from the eHYD platform (<uri>https://ehyd.gv.at</uri>, last access: 19 January 2026) are available with a time delay of only 2 h, making the integration of recent discharge observations operationally feasible for real-time forecasting applications.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e2662">Cumulative Density Function of Nash Sutcliffe Efficiency values for experiments integrating past discharge in the hindcast phase. The experiments include the baselines (grey), Encoder- Decoder LSTM (purple) and Sequential Forecast LSTM (orange). Solid lines depict the experiments without discharge, while the dashed lines show experiments with discharge in the hindcast phase. The arrows represent the prediction accuracy increase as median <inline-formula><mml:math id="M67" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE when integrating discharge in the hindcast phase of the models.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f07.png"/>

        </fig>

      <p id="d2e2678">The integration of past discharge observations demonstrates performance improvements across both tested architectures. For the Encoder–Decoder LSTM, incorporating discharge data yields significant gains, with the median NSE increasing from 0.57 to 0.66. The improvement is particularly pronounced in the lower percentiles, with the 10th percentile rising from 0.21 to 0.37, indicating considerably better performance for poorly performing basins. The 75th and 90th percentiles also show notable improvements (0.65 to 0.77 and 0.73 to 0.83, respectively), while the variability decreases slightly (standard deviation from 8.03 to 6.68).</p>
      <p id="d2e2682">When discharge is integrated, the Sequential Forecast LSTM achieves a median NSE of 0.71, compared to 0.63 without discharge. The absolute improvement (<inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mn mathvariant="normal">0.08</mml:mn></mml:mrow></mml:math></inline-formula>) is marginally smaller than for the Encoder–Decoder LSTM (<inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mn mathvariant="normal">0.09</mml:mn></mml:mrow></mml:math></inline-formula>), but the Sequential Forecast LSTM reaches a higher absolute performance level, surpassing even the reanalysis baseline of 0.69. The 10th percentile shows substantial improvement (0.35 to 0.42), while the upper percentiles reach 0.81 and 0.88 for the 75th and 90th percentiles, also exceeding the reanalysis baseline performance in these ranges.</p>
      <p id="d2e2705">The performance gains from discharge integration are not surprising given that near-real-time discharge observations constrain the current hydrological state through autoregressive inputs (Nearing et al., 2022), providing the model with direct information about catchment conditions at the time of forecast initialization that meteorological inputs alone cannot capture. These gains therefore reflect a combination of effects: beyond reducing meteorological forecast-induced bias, improved hydrological state estimation directly contributes to forecast skill regardless of meteorological input quality.</p>
      <p id="d2e2708">Additional transfer learning experiments were conducted to investigate whether discharge information learned during pre-training could be effectively transferred to discharge-free operational scenarios. These experiments included domain adaption configurations where discharge was incorporated during the pre-training phase but excluded during fine-tuning, as well as setups using the Sequential Forecast LSTM with discharge in the source domain and without discharge in the target domain. All transfer learning approaches consistently resulted in performance deterioration and failed to yield improvements that would justify the computational overhead of the transfer learning process. While these experiments demonstrated that information extraction from LSTM cell states is feasible, the learned discharge-related representations could not adequately compensate for the absence of direct discharge observations during operational forecasting. The experiments with discharge incorporated directly into the training data consistently outperformed all transfer learning alternatives, indicating that real-time discharge integration provides irreplaceable benefits that cannot be effectively substituted through knowledge transfer mechanisms.</p>
</sec>
<sec id="Ch1.S3.SS5">
  <label>3.5</label><title>The role of embedding complexity</title>
      <p id="d2e2719">During experimenting we experienced that the embedding complexity plays an important role in reducing the forecast induced bias in runoff predictions. This circumstance has led us to create this experimental setup, in which a simple (linear) and a complex (non-linear) embedding were generated for all architectures used in the previous experimental settings. In brief, a similar pattern can be observed for all architectures except transfer learning: the more complex the input embedding, the lower the prediction performance.</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e2724">Cumulative Density Function of Nash Sutcliffe Efficiency values for experiments with different embedding complexities. The experiments include the baselines (grey), selective transfer learning with only finetuning the embedding weights while the LSTM was freezed (TL EmbeddingNet), Encoder–Decoder LSTM without discharge (purple), Encoder–Decoder LSTM with discharge (pink), Sequential Forecast LSTM without discharge (orange) and Sequential Forecast LSTM with discharge (yellow). Solid lines depict the experiments without simple linear embedding, while the dashed lines show experiments with complex non- linear embedding networks.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f08.png"/>

        </fig>

      <p id="d2e2733">While the transfer learning experiments did not achieve the high NSE values of the best-performing forecasting approach (Sequential LSTM), notable relationships emerged between embedding complexity and prediction accuracy. In the baseline, encoder–decoder and sequential LSTM experiments, simple linear embedding networks produced slightly higher simulation performance compared to more complex embedding architectures. This pattern reversed in transfer learning experiments, where increased embedding complexity led to improved prediction results. The transfer learning procedure consisted of transferring weights from an LSTM model pre-trained on all available reanalysis data (<inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="script">D</mml:mi><mml:mi mathvariant="normal">RA</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) to a new network architecture with modified input embeddings, where only the embedding weights underwent retraining while the remaining LSTM parameters remained frozen.</p>
      <p id="d2e2748">The minimal performance improvement of simple embeddings in non-transfer learning experiments can be attributed to the tendency of complex embedding networks to overfit when trained on the available data. The number of trainable parameters in complex embeddings exceeds the optimal ratio relative to available training data, causing the model to learn specific noise patterns in the training data rather than generalizable hydrological patterns. This overfitting tendency is further amplified by the absence of dropout regularization within the embedding networks, which was set to 0 across all experiments. Additionally, the use of tanh activations in both the embedding networks and the LSTM may compound saturation effects (Acuña Espinoza et al., 2025), potentially contributing to the performance degradation observed with more complex embeddings. This means simpler embeddings capture the relevant input patterns more effectively without introducing unnecessary model complexity comprising generalization.</p>
      <p id="d2e2751">The improved results with increased embedding complexity in transfer learning can be explained by the requirement for flexible, non-linear transformations necessary for effective domain adaptation. Since only the embedding weights are trained while all other network parameters remain frozen, the embedding layer must perform all adaptation work between the target domain and the representations pre-trained on reanalysis data. More complex architectures can perform richer and more domain-specific feature extraction, compensating for the discrepancy between source and target domains through more expressive input transformations. The contrasting performance patterns between transfer learning and standard training suggest that the optimal embedding complexity depends on whether the model parameters are trained from scratch or adapted from pre-trained weights, though the precise mechanisms underlying this relationship warrant further investigation.</p>
</sec>
<sec id="Ch1.S3.SS6">
  <label>3.6</label><title>Physiographic Controls on Model Skill</title>
      <p id="d2e2762">While the previous sections demonstrated the overall effectiveness of the proposed bias correction strategies, a key operational question remains: <italic>does model skill improve uniformly across catchment types, or do specific basin characteristics such as aridity, elevation, land cover, or precipitation regime determine which catchments benefit most from the proposed strategies?</italic></p>
      <p id="d2e2767">Inspired by the benchmark framework of Seibert et al. (2018), who argue that model performance should always be evaluated relative to meaningful reference points, we compute the per-basin <inline-formula><mml:math id="M71" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE between each experimental configuration from Sect. 3.2 and 3.4 to the Baseline Forecast. Figure 9 shows these <inline-formula><mml:math id="M72" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE values with basins sorted by their mean <inline-formula><mml:math id="M73" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE across all experiments in ascending order, providing a clear picture of which basins consistently improve, remain unchanged, or degrade in forecast skill across the proposed strategies.</p>

      <fig id="F9" specific-use="star"><label>Figure 9</label><caption><p id="d2e2793">Per-basin skill difference (<inline-formula><mml:math id="M74" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE) between each model architecture and the Forecast Baseline across 451 gauged catchments of the LamaH-CE dataset. Each row corresponds to one experiment and each column to one catchment. <inline-formula><mml:math id="M75" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE is defined as the difference between the basin NSE values from an architecture (rows) and the Baseline Forecast, where positive values (red) indicate that the architecture outperforms the baseline and negative values (blue) indicate degraded performance. Basins are sorted left to right by their mean <inline-formula><mml:math id="M76" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE across all experiments (ascending), and experiments are sorted top to bottom by their median <inline-formula><mml:math id="M77" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE (descending). The colour scale is symmetric and bounded by the 5th and 95th percentiles of the pooled <inline-formula><mml:math id="M78" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE distribution. The bar chart on the right shows the fraction of catchments with improved (red) and degraded (blue) performance relative to the baseline for each experiment.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f09.png"/>

        </fig>

      <p id="d2e2838">The results in Fig. 9 reveal that the Sequential Forecast LSTM with discharge integration achieves positive <inline-formula><mml:math id="M79" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE values across the vast majority, with 445 of the 451 basins, confirming that nearly all basins benefit from this configuration relative to the Baseline Forecast. A small number of isolated basins show exceptionally large positive <inline-formula><mml:math id="M80" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE values, visible as dark red columns, which are particularly pronounced in the discharge-augmented configurations and likely reflect basins where near real-time discharge observations provide especially strong anchoring of the initial hydrological state. The Baseline Reanalysis shows a very similar basin-level improvement pattern, consistent with its role as the upper performance reference.</p>
      <p id="d2e2855">A clear and consistent architectural difference emerges between the Sequential Forecast LSTM and the Encoder–Decoder LSTM. The Sequential Forecast LSTM without discharge integration already improves 435 out of 451 basins, while the Encoder–Decoder LSTM without discharge improves only 361 out of 451 basins and actively degrades 90, nearly 5 times as many as the Sequential Forecast LSTM. Even with discharge integration, the Encoder–Decoder LSTM degrades 28 basins compared to only 5 for the Sequential Forecast LSTM. This persistent instability is consistent with the higher standard deviations reported for the Encoder–Decoder architecture in Sect. 3.2 and suggests that the fixed handoff network introduces fragility for specific basin conditions that the Sequential LSTM's continuous state evolution avoids.</p>
      <p id="d2e2858">The transfer learning configurations show the most heterogeneous basin-level patterns. TL EmbeddingNet and TL AllWeights improve only 275 and 273 out of 451 basins respectively, while degrading the rest, indicating that both transfer learning strategies fail to provide consistent improvements across the basin population and cannot be considered reliable bias correction strategies across diverse catchment conditions.</p>
      <p id="d2e2861">Since we now know which catchments benefit most from each architecture in terms of increased forecasting skill, we further investigate which catchment characteristics determine where the largest improvements occur. To this end, we focus exclusively on basins with positive <inline-formula><mml:math id="M81" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE values (red columns in Fig. 9) – that is, basins where the proposed strategy outperforms the Baseline Forecast. Basins with negative <inline-formula><mml:math id="M82" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE are excluded from this analysis. Figure 10 shows the Spearman correlation between the per-basin <inline-formula><mml:math id="M83" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE and the 33 static catchment attributes across all experimental configurations for this subset of basins. Positive values indicate that catchments with larger values of a given attribute tend to benefit more from the proposed strategy, while negative values indicate smaller improvements.</p>

      <fig id="F10" specific-use="star"><label>Figure 10</label><caption><p id="d2e2887">Spearman correlation of experiments with improved model skill compared to the Forecast Baseline, denoted as <inline-formula><mml:math id="M84" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula>. Each row corresponds to one experiment and each column to a static catchment attribute used in model training. The Asterisk shows the non significant attributes (<inline-formula><mml:math id="M86" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>).</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f10.png"/>

        </fig>

      <p id="d2e2926">The Spearman correlation between per-basin <inline-formula><mml:math id="M87" display="inline"><mml:mi mathvariant="normal">Δ</mml:mi></mml:math></inline-formula>NSE and static catchment attributes reveals a consistent and striking pattern across all experimental configurations. Aridity index (arid_1) and mean precipitation (p_mean) emerge as the dominant controls on the magnitude of improvement, with aridity showing positive correlations across all experiments (ranging from <inline-formula><mml:math id="M88" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.31</mml:mn></mml:mrow></mml:math></inline-formula> for TL AllWeights to <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.50</mml:mn></mml:mrow></mml:math></inline-formula> for the Sequential Forecast LSTM with discharge) and mean precipitation showing consistently negative correlations of similar magnitude (<inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.31</mml:mn></mml:mrow></mml:math></inline-formula> to <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:mi>r</mml:mi><mml:mo>=</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.50</mml:mn></mml:mrow></mml:math></inline-formula>). This indicates that drier catchments benefit more from the proposed bias correction strategies in absolute terms, while wetter catchments show smaller improvements over the Baseline Forecast.</p>
      <p id="d2e2988">When combined with the Wasserstein distance analysis of Sect. 3.1.1, these results reveal a physically meaningful asymmetry. The attributes showing larger meteorological domain shift between reanalysis and forecast inputs, including mean precipitation, mean slope, fraction of snow and aridity, are precisely the same attributes associated with smaller improvements in forecast skill across all experimental configurations. Conversely, the basins associated with smaller meteorological domain shift, characterized by higher aridity and more frequent dry days, are the same associated with larger gains in forecast skill relative to the Baseline Forecast (see Fig. 11). We explain this paradox by the nature of the hydrological signal in different catchment types: alpine, wet, and snow-dominated catchments experience the largest distributional differences between reanalysis and ECMWF-HRES inputs, but their hydrology is dominated by a strong and seasonally predictable snowmelt signal. This signal is captured and retained in the LSTM cell state through its recurrent processing of the 365 d input sequence, enabling the model to compensate for the large meteorological domain shift through its long-term memory mechanism even when forecast inputs are strongly biased. This interpretation is consistent with the findings of Kratzert et al. (2019a), who demonstrated that LSTM memory cells internally learn to represent snow storage dynamics even without snow being explicitly provided as a model input, confirming the ability of LSTMs to encode physically meaningful hydrological states in their cell states. As a result, even the Baseline Forecast achieves relatively good performance in alpine catchments, leaving less room for improvement from the proposed strategies. In contrast, drier and more episodic catchments experience smaller meteorological domain shifts but have weaker and more nonlinear rainfall-runoff relationships that are more sensitive to input quality, leading to lower Baseline Forecast performance and consequently larger absolute improvements when different model architectings or training strategies are applied.</p>

      <fig id="F11" specific-use="star"><label>Figure 11</label><caption><p id="d2e2993">Maps of the 1-Wasserstein distance between Reanalysis and ECMWF-HRES Forecast distributions for five meteorological variables (Precipitation, Temperature, Actual Evapotranspiration, Solar Radiation and Dewpoint Temperature) across the 451 LamaH basins during the test period (2014–2017), shown as basin polygons colored by Wasserstein distance. Gauge locations are overlaid as circles colored by the per-basin NSE of the Baseline Forecast experiment, providing a direct spatial comparison between the magnitude of meteorological domain shift and the model skill achieved without any bias correction strategy. Larger Wasserstein distances indicate greater distributional differences between the two data products at a given basin.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f11.png"/>

        </fig>

      <p id="d2e3002">This finding has important operational implications. While the proposed strategies consistently outperform the Baseline Forecast across the vast majority of basins as shown in the CDFs, the largest absolute gains are concentrated in arid catchments where the hydrological response is more complex and the baseline performance is lowest, rather than in the topographically complex alpine catchments where the meteorological input bias is largest.</p>
</sec>
<sec id="Ch1.S3.SS7">
  <label>3.7</label><title>Limitations and Future Directions</title>
      <p id="d2e3013">This study is subject to several limitations. The experiments were conducted on the Extended LamaH-CE dataset, which is restricted to Central Europe; the transferability of the results to other hydroclimatic regions remains to be tested. While discharge observations were shown to substantially improve performance, they are not error-free, and their availability cannot be guaranteed in ungauged or poorly monitored catchments. The focus on a 1 d lead time restricts the conclusions that can be drawn regarding forecast quality decay with increasing lead time, and extension to multi-day-ahead forecasting is the subject of a forthcoming study.</p>
      <p id="d2e3016">In addition, we only evaluated existing LSTM-based architectures and a modeling setup with spatially lumped data and daily maximum discharges. Applying these approaches in fully distributed settings with higher temporal resolution forecasts may yield different, potentially more pronounced results.</p>
      <p id="d2e3019">The use of ECMWF-HRES as the only forecast dataset was driven by data availability – to our knowledge, ECMWF-HRES is the only NWP system for which sufficiently long archives of forecast data are openly available to cover the experimental period used in this study. From a theoretical standpoint, however, the approach is not inherently tied to any specific NWP system. A key advantage of using discharge as the target variable is that the model trains on an integrated catchment signal that implicitly captures the combined effect of all meteorological inputs, regardless of their source, meaning the training strategy can in principle be applied to any NWP system by replacing or supplementing the forecast inputs accordingly. Drawing on the findings of Kratzert et al. (2021), who demonstrated that combining multiple meteorological data sources improves LSTM simulation accuracy in a reanalysis setting, we further hypothesize that combining forecasts from multiple NWP systems could yield similar benefits in a forecasting context, enabling the model to learn source-specific bias patterns and potentially improving robustness across different forecast systems</p>
      <p id="d2e3022">The integration of reanalysis data in the hindcast phase, while useful for training, is also constrained by data latency, which is often critical for operational use, although future improvements in near–real-time reanalysis may alleviate this. However, this limitation depends strongly on the application scale. At national scales, real-time analysis, nowcast, or station-based observed data can serve as the hindcast forcing, bypassing reanalysis latency entirely. Products such as INCA from GeoSphere Austria (Haiden et al., 2011) are a good example of operationally available high-resolution analysis data that could fill this role. In large-sample settings such as ours, however, station-based data could introduce further practical challenges, as the number of stations per basin varies greatly, a large basin may contain dozens of stations while a small one may have only one, complicating consistent spatial aggregation and resulting in varying input dimensions across basins. While input embeddings as used in this study could potentially help map these varying input dimensions into a shared latent space, the implications for model performance and stability in such a setting remain an open question.</p>
      <p id="d2e3026">Addressing these challenges – testing transferability, reducing dependence on delayed or unavailable inputs, and extending both the spatial and temporal resolution of experiments – will be essential to further investigate and advance robust, operationally viable solutions for flood forecasting.</p>
</sec>
</sec>
<sec id="Ch1.S4" sec-type="conclusions">
  <label>4</label><title>Conclusion and Outlook</title>
      <p id="d2e3038">The operational deployment of deep learning based flood forecasting models faces fundamental challenges when transitioning from high-quality reanalysis to meteorological forecast data, with domain shift between these data sources leading to model performance degradation. While previous approaches have focused on bias-correcting meteorological inputs through statistical methods (Lenderink et al., 2007) or machine learning techniques applied to precipitation forecasts (Ko et al., 2020; Zhang et al., 2020), these methods rely on comparing forecasts with meteorological observations that themselves contain uncertainties (Bárdossy et al., 2022). This study addressed the challenge through systematic evaluation of Long Short-Term Memory architectures and training techniques that learn bias correction directly from the more reliable discharge observations, following the paradigm suggested by Kirchner (2009) of using river discharge as the primary constraint. Our experiments across 451 Central European catchments demonstrated that appropriate neural network designs can transform the domain shift problem from a major obstacle into a learnable pattern correction task. Sequential Forecast LSTM architectures, when combining meteorological hindcast data with past discharge observations, provided the most effective framework for mitigating forecast-induced biases. This configuration achieved a median NSE of 0.71, surpassing even the reanalysis baseline simulation and establishing discharge integration, if data are available in near real time as in the LamaH domain, as a critical component for operational forecast accuracy.</p>
      <p id="d2e3041">To quantify the bias propagation caused by the domain shift, we conducted cross-domain evaluation revealing performance deterioration when reanalysis-trained models were applied to forecast inputs. In this setting, the median Nash-Sutcliffe Efficiency decreased from 0.58 to 0.33, representing a 0.25 reduction in model skill. This performance degradation stems from fundamental differences in data distributions between reanalysis and forecast datasets, violating the assumption of identically distributed training and testing data that underlies machine learning model generalization (Goodfellow et al., 2016). Analysis of the spatial patterns of this input data domain shift, quantified using the 1-Wasserstein distance across the 451 basins, reveals that topographically complex, high-elevation, and snow-dominated catchments experience the largest distributional differences between reanalysis and ECMWF-HRES inputs, potentially driven by the limited ability of numerical weather prediction models to resolve orographic effects at their operational resolution (Haiden et al., 2024; Lavers et al., 2021).</p>
      <p id="d2e3044">Among the tested neural network architectures, the Sequential Forecast LSTM demonstrated superior performance for operational forecasting applications. This architecture achieved a median NSE of 0.63 with notable stability (standard deviation of 0.52) and maintained reasonable performance across all percentiles. The sequential processing approach enables gradual correction of forecast biases through continuous state evolution from hindcast to forecast phases, preserving temporal patterns without the disruption introduced by fixed handoff networks in Encoder–Decoder architectures. Transfer learning approaches, despite theoretical advantages for domain adaptation, achieved only modest improvements (NSE 0.44), while the incorporation of archived forecasts in training failed to provide consistent benefits and often increased model instability. The relationship between embedding complexity and performance varied systematically: simpler embeddings performed better in standard training contexts, while complex embeddings showed advantages only in transfer learning scenarios where flexible input transformations were required for domain adaptation.</p>
      <p id="d2e3047">These findings carry important implications for operational flood forecasting system design. The high performance of Sequential Forecast LSTM architectures indicates that operational systems should prioritize continuous state transfer mechanisms maintaining temporal dependencies across the hindcast-forecast phases rather than treating these phases as disconnected processes. The improvements from discharge integration are consistent with the operational capabilities of many monitored systems, such as hydro power plants or the Austrian eHYD platform, where observations are available with a minimal latency, making real-time integration feasible. The ability to learn bias correction patterns directly from the combined meteorological-hydrological data space eliminates the need for separate pre-processing steps, whether meteorological bias correction (Hess, 2020; Han et al., 2021) or streamflow post-processing methods like the approach of Hunt et al. (2022), reducing computational overhead and possibly potential error propagation. Basin-level analysis of forecast skill improvements, measured as the difference in NSE between each tested architecture and the Baseline Forecast, reveals that the proposed strategies do not improve performance equally across catchment types. While the Sequential Forecast LSTM with discharge integration outperforms the Baseline Forecast in 445 out of 451 basins, the magnitude of improvement varies systematically with catchment characteristics. The largest absolute gains are concentrated in arid and precipitation-limited catchments, where weaker and more nonlinear rainfall-runoff relationships make model skill more sensitive to input quality. In contrast, alpine and snow-dominated catchments, despite experiencing the largest meteorological domain shift, show smaller improvements – a phenomenon we explain with the LSTM cell state capturing and retaining the strong seasonal snowmelt signal through its long-term memory mechanism, enabling effective compensation for forecast input bias even without explicit bias correction. Future developments should focus on adaptive architectures that can dynamically leverage discharge observations when available while maintaining robust performance in ungauged settings through reanalysis-only hindcast processing. Such unified frameworks would enable seamless deployment across both gauged and ungauged basins within the same operational system, automatically adjusting to data availability in real-time. However, sensor failures should also be taken into account here, and the training methods proposed by Gauch et al. (2025) should be applied. While our experiments were limited to ECMWF HRES archived forecasts due to data availability, combining multiple forecasts from different sources could also be promising as it could capture forecast uncertainty more comprehensively and enable the model to learn source-specific bias patterns, as with the integration of multiple meteorological data in a simulation setting (Kratzert et al., 2021).  The Sequential Forecast LSTM's bias correction capabilities at 24 h lead times provide a strong foundation for multi-day forecasting applications, where learning lead-time dependent bias patterns could improve medium-range flood predictions that are crucial for early warning systems and emergency preparedness. The demonstrated ability of LSTM architectures and training techniques to transform the domain shift challenge into a learnable bias correction problem, combined with increasing availability of real-time hydrological observations, establishes a pathway toward operational flood forecasting systems that can maintain predictive skill despite the inherent uncertainties coming from numerical weather predictions. Beyond the methodological advances presented here, the spatially heterogeneous patterns of domain shift and forecast skill improvement identified across the 451 study basins open a promising avenue for future research: understanding why forecast-induced bias manifests differently across catchment types in terms of its temporal dynamics, physiographic controls, and physical mechanisms. This represents the natural next step toward a comprehensive theory of forecast uncertainty in data-driven hydrological modelling.</p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <label>Appendix A</label><title>Dynamic and Static Forcings of the Models</title>

<table-wrap id="TA1a"><label>Table A1</label><caption><p id="d2e3067">Meteorological Variables in the Extended LamaH-CE data set used for the conducted experiments. All reanalysis variables are used as dynamic inputs across all experiments, with two exceptions: the Cross Domain evaluation, where only the five variables with a direct equivalent in the ECMWF-HRES forecast data are used (see Table 1) to ensure an identical input feature space between training and inference; and the Baseline Forecast experiment, where only the five ECMWF-HRES forecast variables are used as dynamic inputs.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="4cm"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Variable</oasis:entry>
         <oasis:entry colname="col2" align="left">Description</oasis:entry>
         <oasis:entry colname="col3">Unit</oasis:entry>
         <oasis:entry colname="col4">Source Product</oasis:entry>
         <oasis:entry colname="col5">Source</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_2m_temp_max</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface max air temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_2m_temp_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface mean air temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_2m_temp_min</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface min air temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_2m_dp_temp_max</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface max dewpoint temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_2m_dp_temp_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface mean dewpoint temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_2m_dp_temp_min</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface min dewpoint temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_10m_wind_u</oasis:entry>
         <oasis:entry colname="col2" align="left">Eastwards wind speed 10 m above earth surface</oasis:entry>
         <oasis:entry colname="col3">m s<sup>−1</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_10m_wind_v</oasis:entry>
         <oasis:entry colname="col2" align="left">Northwards wind speed 10 m above earth surface</oasis:entry>
         <oasis:entry colname="col3">m s<sup>−1</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_fcst_alb</oasis:entry>
         <oasis:entry colname="col2" align="left">Forecast albedo</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_lai_high_veg</oasis:entry>
         <oasis:entry colname="col2" align="left">Leaf Area Index for high vegetation type</oasis:entry>
         <oasis:entry colname="col3">m<sup>2</sup> m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_lai_low_veg</oasis:entry>
         <oasis:entry colname="col2" align="left">Leaf Area Index for low vegetation type</oasis:entry>
         <oasis:entry colname="col3">m<sup>2</sup> m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_swe</oasis:entry>
         <oasis:entry colname="col2" align="left">Snow Water Equivalent</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_surf_net_solar_rad_max</oasis:entry>
         <oasis:entry colname="col2" align="left">Max amount of solar radiation  reaching the Earth's surface minus the amount reflected by the Earth's surface</oasis:entry>
         <oasis:entry colname="col3">W m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_surf_net_solar_rad_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean amount of solar radiation  reaching the Earth's surface minus the amount reflected by the Earth's surface</oasis:entry>
         <oasis:entry colname="col3">W m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_surf_net_therm_rad_max</oasis:entry>
         <oasis:entry colname="col2" align="left">Maximum net thermal radiation at the Earth's surface;</oasis:entry>
         <oasis:entry colname="col3">W m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_surf_net_therm_rad_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean net thermal radiation at the Earth's surface;</oasis:entry>
         <oasis:entry colname="col3">W m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_surf_press</oasis:entry>
         <oasis:entry colname="col2" align="left">Surface pressure</oasis:entry>
         <oasis:entry colname="col3">Pa</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_total_et</oasis:entry>
         <oasis:entry colname="col2" align="left">Total evapotranspiration</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ERA5L_prec</oasis:entry>
         <oasis:entry colname="col2" align="left">Total precipitation</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="TA1b"><label>Table A1</label><caption><p id="d2e3566">Continued.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="5">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="4cm"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Variable</oasis:entry>
         <oasis:entry colname="col2" align="left">Description</oasis:entry>
         <oasis:entry colname="col3">Unit</oasis:entry>
         <oasis:entry colname="col4">Source Product</oasis:entry>
         <oasis:entry colname="col5">Source</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_volsw_123</oasis:entry>
         <oasis:entry colname="col2" align="left">Fraction of water from 0 to 100 cm depth (topsoil)</oasis:entry>
         <oasis:entry colname="col3">m<sup>3</sup> m<sup>−3</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. (2021)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ERA5L_volsw_4</oasis:entry>
         <oasis:entry colname="col2" align="left">Fraction of water from 100 to 289 cm depth (subsoil)</oasis:entry>
         <oasis:entry colname="col3">m<sup>3</sup> m<sup>−3</sup></oasis:entry>
         <oasis:entry colname="col4">ERA5Land</oasis:entry>
         <oasis:entry colname="col5">Muñoz-Sabater et al. 2021</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">EOBS_tg</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface mean daily air temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">E-OBS</oasis:entry>
         <oasis:entry colname="col5">Cornes et al. (2018)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">EOBS_tn</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface min daily air temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">E-OBS</oasis:entry>
         <oasis:entry colname="col5">Cornes et al. (2018)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">EOBS_tx</oasis:entry>
         <oasis:entry colname="col2" align="left">2 m above earth surface max daily air temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">E-OBS</oasis:entry>
         <oasis:entry colname="col5">Cornes et al. (2018)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">EOBS_rr</oasis:entry>
         <oasis:entry colname="col2" align="left">Total precipitation</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">E-OBS</oasis:entry>
         <oasis:entry colname="col5">Cornes et al. (2018)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">EOBS_pp</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean sea level pressure</oasis:entry>
         <oasis:entry colname="col3">hPa</oasis:entry>
         <oasis:entry colname="col4">E-OBS</oasis:entry>
         <oasis:entry colname="col5">Cornes et al. (2018)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">EOBS_fg</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean wind speed at 10 m height</oasis:entry>
         <oasis:entry colname="col3">m s<sup>−1</sup></oasis:entry>
         <oasis:entry colname="col4">E-OBS</oasis:entry>
         <oasis:entry colname="col5">Cornes et al. (2018)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">EOBS_qq</oasis:entry>
         <oasis:entry colname="col2" align="left">Solar radiation at earth's surface</oasis:entry>
         <oasis:entry colname="col3">W m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">E-OBS</oasis:entry>
         <oasis:entry colname="col5">Cornes et al. (2018)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">MSWEP_RR</oasis:entry>
         <oasis:entry colname="col2" align="left">Total precipitation</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">MSWEP</oasis:entry>
         <oasis:entry colname="col5">Beck et al. (2019)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">GLEAM_ETA</oasis:entry>
         <oasis:entry colname="col2" align="left">Actual evapotranspiration</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">GLEAM</oasis:entry>
         <oasis:entry colname="col5">Miralles et al. (2011)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">GLEAM_ETP</oasis:entry>
         <oasis:entry colname="col2" align="left">Potential evapotranspiration</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">GLEAM</oasis:entry>
         <oasis:entry colname="col5">Miralles et al. (2011)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ECMWF_t2m</oasis:entry>
         <oasis:entry colname="col2" align="left">Forecasted 2m above earth surface mean air temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">ECMWF-HRES</oasis:entry>
         <oasis:entry colname="col5">ECMWF (2025)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ECMWF_d2m</oasis:entry>
         <oasis:entry colname="col2" align="left">Forecasted 2 m above earth surface mean dewpoint temperature</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4">ECMWF-HRES</oasis:entry>
         <oasis:entry colname="col5">ECMWF (2025)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ECMWF_ssrd</oasis:entry>
         <oasis:entry colname="col2" align="left">Forecasted solar radiation at earth's surface</oasis:entry>
         <oasis:entry colname="col3">W m<sup>−2</sup></oasis:entry>
         <oasis:entry colname="col4">ECMWF-HRES</oasis:entry>
         <oasis:entry colname="col5">ECMWF (2025)</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ECMWF_tp</oasis:entry>
         <oasis:entry colname="col2" align="left">Forecasted total precipitation</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">ECMWF-HRES</oasis:entry>
         <oasis:entry colname="col5">ECMWF (2025)</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ECMWF_e</oasis:entry>
         <oasis:entry colname="col2" align="left">Forecasted total actual evapotranspiration</oasis:entry>
         <oasis:entry colname="col3">mm</oasis:entry>
         <oasis:entry colname="col4">ECMWF-HRES</oasis:entry>
         <oasis:entry colname="col5">ECMWF (2025)</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<table-wrap id="TA2"><label>Table A2</label><caption><p id="d2e3996">Static catchment attributes from the Extended LamaH-CE data set used for the conducted experiments.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="13cm"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Attribute</oasis:entry>
         <oasis:entry colname="col2" align="left">Description</oasis:entry>
         <oasis:entry colname="col3">Unit</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">area_calc</oasis:entry>
         <oasis:entry colname="col2" align="left">Calculated basin area</oasis:entry>
         <oasis:entry colname="col3">km<sup>2</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">elev_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean catchment elevation</oasis:entry>
         <oasis:entry colname="col3">m a.s.l.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">elev_med</oasis:entry>
         <oasis:entry colname="col2" align="left">Median catchment elevation</oasis:entry>
         <oasis:entry colname="col3">m a.s.l.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">elev_std</oasis:entry>
         <oasis:entry colname="col2" align="left">Standard deviation of elevation in catchment</oasis:entry>
         <oasis:entry colname="col3">m a.s.l.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">elev_ran</oasis:entry>
         <oasis:entry colname="col2" align="left">Range of catchment elevation (max–min elev.)</oasis:entry>
         <oasis:entry colname="col3">m a.s.l.</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">slope_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean catchment slope</oasis:entry>
         <oasis:entry colname="col3">m km<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">mvert_dist</oasis:entry>
         <oasis:entry colname="col2" align="left">Horizontal distance from the farthest point of the catchment to the corresponding gauge (length axis)</oasis:entry>
         <oasis:entry colname="col3">km<sup>2</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">mvert_ang</oasis:entry>
         <oasis:entry colname="col2" align="left">Angle between the north direction and connection from farthest point of catchment to the corresponding gauge (length axis)</oasis:entry>
         <oasis:entry colname="col3">degree</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">elon_ratio</oasis:entry>
         <oasis:entry colname="col2" align="left">Elongation ratio between the diameter <inline-formula><mml:math id="M112" display="inline"><mml:mi>D</mml:mi></mml:math></inline-formula> of an equicalent circle and the area of the catchment area to <inline-formula><mml:math id="M113" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>st length <inline-formula><mml:math id="M114" display="inline"><mml:mi>L</mml:mi></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">strm_dens</oasis:entry>
         <oasis:entry colname="col2" align="left">Stream density</oasis:entry>
         <oasis:entry colname="col3">km km<sup>−2</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">p_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean daily precipitation</oasis:entry>
         <oasis:entry colname="col3">mm d<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">et0_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean daily reference evapotranspiration</oasis:entry>
         <oasis:entry colname="col3">mm d<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">eta_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean daily total evapotranspiration</oasis:entry>
         <oasis:entry colname="col3">mm d<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">arid_1</oasis:entry>
         <oasis:entry colname="col2" align="left">Aridity, computed as the ratio of mean et0_mean and p_pean</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">arid_2</oasis:entry>
         <oasis:entry colname="col2" align="left">Reciprocal value of aridity index</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">p_season</oasis:entry>
         <oasis:entry colname="col2" align="left">Seasonality and timing of precipitation (estimated using sine curves) to represent the annual precipitation cycles</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">frac_snow</oasis:entry>
         <oasis:entry colname="col2" align="left">Fraction of precipitation falling as snow</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">hi_prec_fr</oasis:entry>
         <oasis:entry colname="col2" align="left">Frequency of high-precipitation days</oasis:entry>
         <oasis:entry colname="col3">d yr<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">hi_prec_du</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean duration of high-precipitation events</oasis:entry>
         <oasis:entry colname="col3">day</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">lo_prec_fr</oasis:entry>
         <oasis:entry colname="col2" align="left">Frequency of dry days</oasis:entry>
         <oasis:entry colname="col3">d yr<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">lo_prec_du</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean duration of dry periods</oasis:entry>
         <oasis:entry colname="col3">day</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">lc_dom</oasis:entry>
         <oasis:entry colname="col2" align="left">Three-digit short code of dominant land cover class</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">agr_fra</oasis:entry>
         <oasis:entry colname="col2" align="left">Fraction of agricultural areas</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">bare_fra</oasis:entry>
         <oasis:entry colname="col2" align="left">Fraction of bare areas</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">forest_fra</oasis:entry>
         <oasis:entry colname="col2" align="left">Fraction of forest areas</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">lake_fra</oasis:entry>
         <oasis:entry colname="col2" align="left">Fraction of natural or artificial water bodies with all-season water filling</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">urban_fra</oasis:entry>
         <oasis:entry colname="col2" align="left">Fraction of areas mainly occupied by buildings including their connected areas</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">lai_max</oasis:entry>
         <oasis:entry colname="col2" align="left">Maximum monthly mean of one-sided leaf area index</oasis:entry>
         <oasis:entry colname="col3">m<sup>2</sup> m<sup>−2</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">lai_diff</oasis:entry>
         <oasis:entry colname="col2" align="left">Difference between maximum and minimum monthly mean of one-sided leaf area index</oasis:entry>
         <oasis:entry colname="col3">m<sup>2</sup> m<sup>−2</sup></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ndvi_max</oasis:entry>
         <oasis:entry colname="col2" align="left">Maximum monthly mean of NDVI</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">ndvi_min</oasis:entry>
         <oasis:entry colname="col2" align="left">Minimum monthly mean of NDVI</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">gvf_max</oasis:entry>
         <oasis:entry colname="col2" align="left">Maximum monthly mean of the green vegetation fraction</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">gvf_diff</oasis:entry>
         <oasis:entry colname="col2" align="left">Difference between the maximum and minimum monthly mean of the green vegetation fraction</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>


</app>

<app id="App1.Ch1.S2">
  <label>Appendix B</label><title>Variability across input data sets</title>
      <p id="d2e4592">The violin plots in Fig. B1 show the distributions of annually aggregated meteorological variables across all data products and 451 basins. It should be noted that these distributions reflect annual aggregates, and differences at the daily resolution used in the experiments may deviate from the patterns shown here. For temperature and dewpoint temperature, ERA5-Land shows a wider spread than ECMWF-HRES, with ERA5-Land reaching lower minimum values, while the medians differ by less than 0.2 °C between the two products. For precipitation, ERA5-Land reports the highest median annual sum (1223 mm), while E-OBS and MSWEP report notably lower medians (1047 and 1119 mm respectively), with ECMWF-HRES falling in between (1175 mm). The most pronounced differences are found for solar radiation, where ECMWF-HRES reports a median annual sum approximately 11 000 W m<sup>−2</sup> higher than ERA5-Land, and ERA5-Land exhibits a substantially wider and lower-reaching distribution. For evapotranspiration, ERA5-Land and GLEAM show wider spreads with notably lower minimum values compared to ECMWF-HRES, while the medians differ by up to 29 mm annually. These distributional differences across all five variables collectively represent the domain shift that the models must learn to compensate for when transitioning from reanalysis to forecast inputs, as quantified by the 1-Wasserstein distances in Sect. 3.1.1.</p><fig id="FB1"><label>Figure B1</label><caption><p id="d2e4609">Variability across input data sets displayed as violin plots for annually aggregated meteorological variables Temperature, Precipitation, Actual Evapotranspiration, Solar Radiation and Dewpoint Temperature. The violin colours belong always to a certain data product: red- ERA5Land, orange: GLEAM, blue: E-OBS, green: ECMWF-HRES, purple: MSWEP. The dark grey box in the violin shows a boxplot with the bold part depicting the interquartile range and the white line indicating the median.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f12.png"/>

      </fig>


</app>

<app id="App1.Ch1.S3">
  <label>Appendix C</label><title>Evaluation Metrics</title>
      <p id="d2e4630">The model performance was evaluated by using the non basin specific NSE* of Kratzert et al. (2019c) which is based on the Nash and Sutcliffe Efficiency (NSE; Nash and Sutcliffe, 1970). In Eq. (C2), <inline-formula><mml:math id="M126" display="inline"><mml:mi>B</mml:mi></mml:math></inline-formula> is the number of basins, <inline-formula><mml:math id="M127" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> is the number of days per basin <inline-formula><mml:math id="M128" display="inline"><mml:mi>B</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M129" display="inline"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:math></inline-formula> is the predicted value for sample <inline-formula><mml:math id="M130" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the corresponding observed value and <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:mi>s</mml:mi><mml:mo>(</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the training periods standard deviation of the discharge in basin <inline-formula><mml:math id="M133" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula>. For numerical stability, the term <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:mo>∈</mml:mo><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula> is added to the denominator. In comparison to former studies (Kratzert et al., 2019b, c, 2021; Nearing et al., 2024), we conducted no cross validation across spatial units, since we did not focus on ungauged basins. Instead, we employed a temporal split validation approach, where the available time series data for each catchment was divided into training, validation, and testing periods to ensure robust model evaluation.

              <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M135" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="App1.Ch1.S3.E17"><mml:mtd><mml:mtext>C1</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="normal">NSE</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi mathvariant="normal">obs</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi mathvariant="normal">sim</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:munderover><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mi mathvariant="normal">obs</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi mathvariant="normal">obs</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="App1.Ch1.S3.E18"><mml:mtd><mml:mtext>C2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:mi mathvariant="normal">NSE</mml:mi><mml:mo>*</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>B</mml:mi></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>b</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>B</mml:mi></mml:munderover><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>n</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>N</mml:mi></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mo>(</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>n</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:mo>(</mml:mo><mml:mi>s</mml:mi><mml:mo>(</mml:mo><mml:mi>b</mml:mi><mml:mo>)</mml:mo><mml:mo>+</mml:mo><mml:mo>∈</mml:mo><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula></p>
</app>

<app id="App1.Ch1.S4">
  <label>Appendix D</label><title>Hyperparameter Optimization</title>
      <p id="d2e4917">Hyperparameter tuning is a critical component in the bias correction of meteorological forecasting data using Long Short-Term Memory (LSTM) networks, as it directly influences the model's ability to learn complex temporal patterns and correct systematic biases in forecast inputs. Given the nonlinear and dynamic nature of meteorological variables, appropriate selection of hyperparameters such as learning rate, sequence length, number of hidden units, dropout rate, and batch size, is essential to ensure that the LSTM model generalizes well without overfitting to noise or underfitting relevant signals. In the context of bias correction, the model must not only capture historical dependencies in the forecast errors but also effectively differentiate between genuine atmospheric variability and persistent model biases. Without careful tuning, the LSTM may fail to correct biases accurately, particularly under extreme events or seasonal transitions.</p>
      <p id="d2e4920">To this end, we used Bayesian optimization as described by Frazier (2018). This search algorithm fits a Gaussian process to the observed hyperparameter-performance pairs to estimate performance on yet-untested parameter settings. We use the expected improvement as acquisition function, which is used for selecting the next set of hyperparameters to test. This approach efficiently identifies good hyperparameters, especially in large search spaces.</p>
      <p id="d2e4923">For our models, we spanned the search space over the hidden layer sizes (64, 128, 256 units), output dropout rates (0.1, 0.2, 0.3), variance of Gaussian noise applied to the discharge values (0.001, 0.01, 0.1), and batch sizes (64, 128, 256) and limited the number of iterations to 100. To obtain the result for each setting, we trained a model for 30 epochs, an initial learning rate of 0.001 and a cosine annealing schedule (T_max <inline-formula><mml:math id="M136" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 30, <inline-formula><mml:math id="M137" display="inline"><mml:mi mathvariant="italic">η</mml:mi></mml:math></inline-formula>_min <inline-formula><mml:math id="M138" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> <inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:msup><mml:mn mathvariant="normal">10</mml:mn><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:math></inline-formula>), stopping early if the model did not improve the evaluation metric by more than 0.005 in the last 5 epochs. To estimate the performance of the model we used the basin averaged Nash-Sutcliffe efficiency (NSE*) metric and evaluated it on the validation period.</p>
      <p id="d2e4965">NSE* with hyperparameter optimization:

          <disp-formula id="App1.Ch1.S4.E19" content-type="numbered"><label>D1</label><mml:math id="M140" display="block"><mml:mrow><mml:mi mathvariant="italic">λ</mml:mi><mml:mo>=</mml:mo><mml:mi mathvariant="normal">arg</mml:mi><mml:munder><mml:mi mathvariant="normal">max</mml:mi><mml:mo mathvariant="italic">λ</mml:mo></mml:munder><mml:msub><mml:mtext>NSE</mml:mtext><mml:mi mathvariant="normal">median</mml:mi></mml:msub><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi mathvariant="normal">val</mml:mi></mml:msub><mml:mo>;</mml:mo><mml:mi mathvariant="italic">λ</mml:mi></mml:mrow></mml:mfenced></mml:mrow></mml:math></disp-formula>

        where

          <disp-formula id="App1.Ch1.S4.E20" content-type="numbered"><label>D2</label><mml:math id="M141" display="block"><mml:mrow><mml:mi mathvariant="italic">λ</mml:mi><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>h</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>p</mml:mi><mml:mi mathvariant="normal">dropout</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">noise</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi mathvariant="normal">size</mml:mi></mml:msub><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></disp-formula>

        represents the optimal output.</p>
</app>

<app id="App1.Ch1.S5">
  <label>Appendix E</label><title>Long-Short Term Memory Network</title>
      <p id="d2e5049">Long term information is stored in the cell state (<inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), short term information in the hidden state (<inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and the information flow is controlled by the so- called gating mechanisms (Hochreiter and Schmidhuber, 1997). The input gate (<inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:msub><mml:mi>i</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) determines how much of the current input and previous hidden state contributes to updating the cell state. The forget gate (<inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:msub><mml:mi>f</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) regulates which parts of the previous cell state should be retained or discarded, allowing the model to reset its memory when necessary. The output gate (<inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:msub><mml:mi>o</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) defines how much of the updated cell state is exposed to the next time step via the hidden state (Gers et al., 2000). This architecture allows LSTMs to retain relevant information over longer time periods and capture temporal dependencies in input data.

              <disp-formula specific-use="gather" content-type="numbered"><mml:math id="M147" display="block"><mml:mtable displaystyle="true"><mml:mlabeledtr id="App1.Ch1.S5.E21"><mml:mtd><mml:mtext>E1</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>f</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="App1.Ch1.S5.E22"><mml:mtd><mml:mtext>E2</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi mathvariant="normal">tanh</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mover accent="true"><mml:mi>c</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mover accent="true"><mml:mi>c</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover></mml:msub><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mover accent="true"><mml:mi>c</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="App1.Ch1.S5.E23"><mml:mtd><mml:mtext>E3</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="App1.Ch1.S5.E24"><mml:mtd><mml:mtext>E4</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>f</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>i</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>⊙</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="App1.Ch1.S5.E25"><mml:mtd><mml:mtext>E5</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle displaystyle="true" class="stylechange"/><mml:msub><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="italic">σ</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>W</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:msub><mml:mi>x</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>U</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi>b</mml:mi><mml:mi>o</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:mtd></mml:mlabeledtr><mml:mlabeledtr id="App1.Ch1.S5.E26"><mml:mtd><mml:mtext>E6</mml:mtext></mml:mtd><mml:mtd><mml:mrow><mml:mstyle class="stylechange" displaystyle="true"/><mml:msub><mml:mi>h</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="normal">tanh</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>c</mml:mi><mml:mi>t</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>⊙</mml:mo><mml:msub><mml:mi>o</mml:mi><mml:mi>t</mml:mi></mml:msub></mml:mrow></mml:mtd></mml:mlabeledtr></mml:mtable></mml:math></disp-formula></p><fig id="FE1"><label>Figure E1</label><caption><p id="d2e5429">LSTM cell.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f13.png"/>

      </fig>

      <fig id="FE2"><label>Figure E2</label><caption><p id="d2e5442">Sequential Forecast LSTM.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5067/2026/hess-30-5067-2026-f14.png"/>

      </fig>


</app>

<app id="App1.Ch1.S6">
  <label>Appendix F</label><title>Summary Statistics of Model Performance Across All Experimental Configurations</title>

<table-wrap id="TF1"><label>Table F1</label><caption><p id="d2e5468">Summary statistics of Nash-Sutcliffe Efficiency (NSE) values across all 451 basins for each experimental configuration, sorted by median NSE in ascending order. Statistics include the median, mean, standard deviation, minimum, maximum, and the 10th, 25th, 75th, and 90th percentiles. The large differences between mean and median NSE for several configurations, as well as the strongly negative minimum values, reflect the presence of extreme negative outliers in a small number of basins, which disproportionately influence the mean and standard deviation.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="10">
     <oasis:colspec colnum="1" colname="col1" align="justify" colwidth="5cm"/>
     <oasis:colspec colnum="2" colname="col2" align="right"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="right"/>
     <oasis:colspec colnum="10" colname="col10" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Experiment</oasis:entry>
         <oasis:entry colname="col2">Median</oasis:entry>
         <oasis:entry colname="col3">Mean</oasis:entry>
         <oasis:entry colname="col4">SD</oasis:entry>
         <oasis:entry colname="col5">Min</oasis:entry>
         <oasis:entry colname="col6">Max</oasis:entry>
         <oasis:entry colname="col7">p10</oasis:entry>
         <oasis:entry colname="col8">p25</oasis:entry>
         <oasis:entry colname="col9">p75</oasis:entry>
         <oasis:entry colname="col10">p90</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">CrossDomain (Forecast, One Shot)</oasis:entry>
         <oasis:entry colname="col2">0.33</oasis:entry>
         <oasis:entry colname="col3">0.19</oasis:entry>
         <oasis:entry colname="col4">1.1</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">19.05</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.86</oasis:entry>
         <oasis:entry colname="col7"><inline-formula><mml:math id="M149" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.13</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col8">0.16</oasis:entry>
         <oasis:entry colname="col9">0.48</oasis:entry>
         <oasis:entry colname="col10">0.63</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Baseline Forecast</oasis:entry>
         <oasis:entry colname="col2">0.39</oasis:entry>
         <oasis:entry colname="col3">0.25</oasis:entry>
         <oasis:entry colname="col4">1.69</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M150" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">33.29</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.89</oasis:entry>
         <oasis:entry colname="col7">0.02</oasis:entry>
         <oasis:entry colname="col8">0.19</oasis:entry>
         <oasis:entry colname="col9">0.53</oasis:entry>
         <oasis:entry colname="col10">0.68</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">TL EmbeddingNet (SimpleEmbedding)</oasis:entry>
         <oasis:entry colname="col2">0.39</oasis:entry>
         <oasis:entry colname="col3">0.17</oasis:entry>
         <oasis:entry colname="col4">2.96</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">58.49</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.85</oasis:entry>
         <oasis:entry colname="col7">0.08</oasis:entry>
         <oasis:entry colname="col8">0.24</oasis:entry>
         <oasis:entry colname="col9">0.52</oasis:entry>
         <oasis:entry colname="col10">0.64</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">TL AllWeights</oasis:entry>
         <oasis:entry colname="col2">0.41</oasis:entry>
         <oasis:entry colname="col3">0.35</oasis:entry>
         <oasis:entry colname="col4">0.41</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">4.76</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.88</oasis:entry>
         <oasis:entry colname="col7">0.05</oasis:entry>
         <oasis:entry colname="col8">0.25</oasis:entry>
         <oasis:entry colname="col9">0.54</oasis:entry>
         <oasis:entry colname="col10">0.68</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">TL EmbeddingNet (archived Forecast)</oasis:entry>
         <oasis:entry colname="col2">0.41</oasis:entry>
         <oasis:entry colname="col3">0.21</oasis:entry>
         <oasis:entry colname="col4">2.29</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">44.23</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.85</oasis:entry>
         <oasis:entry colname="col7">0.02</oasis:entry>
         <oasis:entry colname="col8">0.25</oasis:entry>
         <oasis:entry colname="col9">0.53</oasis:entry>
         <oasis:entry colname="col10">0.65</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">TL EmbeddingNet (ComplexEmbedding)</oasis:entry>
         <oasis:entry colname="col2">0.44</oasis:entry>
         <oasis:entry colname="col3">0.35</oasis:entry>
         <oasis:entry colname="col4">1.12</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">21.74</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.86</oasis:entry>
         <oasis:entry colname="col7">0.15</oasis:entry>
         <oasis:entry colname="col8">0.29</oasis:entry>
         <oasis:entry colname="col9">0.55</oasis:entry>
         <oasis:entry colname="col10">0.67</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Encoder Decoder LSTM</oasis:entry>
         <oasis:entry colname="col2">0.57</oasis:entry>
         <oasis:entry colname="col3">0.02</oasis:entry>
         <oasis:entry colname="col4">8.03</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">157.91</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.89</oasis:entry>
         <oasis:entry colname="col7">0.21</oasis:entry>
         <oasis:entry colname="col8">0.43</oasis:entry>
         <oasis:entry colname="col9">0.65</oasis:entry>
         <oasis:entry colname="col10">0.73</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Encoder Decoder LSTM (archived Forecast)</oasis:entry>
         <oasis:entry colname="col2">0.57</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M156" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">0.04</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">9.33</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M157" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">183.3</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.9</oasis:entry>
         <oasis:entry colname="col7">0.24</oasis:entry>
         <oasis:entry colname="col8">0.44</oasis:entry>
         <oasis:entry colname="col9">0.69</oasis:entry>
         <oasis:entry colname="col10">0.77</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">CrossDomain (Reanalysis, Pre train)</oasis:entry>
         <oasis:entry colname="col2">0.58</oasis:entry>
         <oasis:entry colname="col3">0.44</oasis:entry>
         <oasis:entry colname="col4">0.87</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">14.7</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.92</oasis:entry>
         <oasis:entry colname="col7">0.13</oasis:entry>
         <oasis:entry colname="col8">0.38</oasis:entry>
         <oasis:entry colname="col9">0.7</oasis:entry>
         <oasis:entry colname="col10">0.79</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Encoder Decoder LSTM (Simple Embedding)</oasis:entry>
         <oasis:entry colname="col2">0.59</oasis:entry>
         <oasis:entry colname="col3">0.35</oasis:entry>
         <oasis:entry colname="col4">3.04</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">60.66</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.91</oasis:entry>
         <oasis:entry colname="col7">0.2</oasis:entry>
         <oasis:entry colname="col8">0.44</oasis:entry>
         <oasis:entry colname="col9">0.69</oasis:entry>
         <oasis:entry colname="col10">0.77</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Sequential Forecast LSTM (Complex Embedding)</oasis:entry>
         <oasis:entry colname="col2">0.61</oasis:entry>
         <oasis:entry colname="col3">0.52</oasis:entry>
         <oasis:entry colname="col4">0.91</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">16.94</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.92</oasis:entry>
         <oasis:entry colname="col7">0.34</oasis:entry>
         <oasis:entry colname="col8">0.48</oasis:entry>
         <oasis:entry colname="col9">0.7</oasis:entry>
         <oasis:entry colname="col10">0.77</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Sequential Forecast LSTM (archived Forecast)</oasis:entry>
         <oasis:entry colname="col2">0.62</oasis:entry>
         <oasis:entry colname="col3">0.19</oasis:entry>
         <oasis:entry colname="col4">6.28</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M161" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">121.96</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.92</oasis:entry>
         <oasis:entry colname="col7">0.31</oasis:entry>
         <oasis:entry colname="col8">0.49</oasis:entry>
         <oasis:entry colname="col9">0.71</oasis:entry>
         <oasis:entry colname="col10">0.78</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Sequential Forecast LSTM</oasis:entry>
         <oasis:entry colname="col2">0.63</oasis:entry>
         <oasis:entry colname="col3">0.57</oasis:entry>
         <oasis:entry colname="col4">0.52</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">8.71</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.9</oasis:entry>
         <oasis:entry colname="col7">0.35</oasis:entry>
         <oasis:entry colname="col8">0.49</oasis:entry>
         <oasis:entry colname="col9">0.72</oasis:entry>
         <oasis:entry colname="col10">0.79</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Encoder Decoder LSTM with <inline-formula><mml:math id="M163" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> (Complex Embedding)</oasis:entry>
         <oasis:entry colname="col2">0.66</oasis:entry>
         <oasis:entry colname="col3">0.21</oasis:entry>
         <oasis:entry colname="col4">6.68</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M164" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">131.42</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.91</oasis:entry>
         <oasis:entry colname="col7">0.37</oasis:entry>
         <oasis:entry colname="col8">0.55</oasis:entry>
         <oasis:entry colname="col9">0.77</oasis:entry>
         <oasis:entry colname="col10">0.83</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Encoder Decoder LSTM with <inline-formula><mml:math id="M165" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> (Simple Embedding)</oasis:entry>
         <oasis:entry colname="col2">0.67</oasis:entry>
         <oasis:entry colname="col3">0.41</oasis:entry>
         <oasis:entry colname="col4">3.83</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M166" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">75.97</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.93</oasis:entry>
         <oasis:entry colname="col7">0.4</oasis:entry>
         <oasis:entry colname="col8">0.57</oasis:entry>
         <oasis:entry colname="col9">0.77</oasis:entry>
         <oasis:entry colname="col10">0.85</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Baseline Forecast &amp; Reanalysis</oasis:entry>
         <oasis:entry colname="col2">0.68</oasis:entry>
         <oasis:entry colname="col3">0.57</oasis:entry>
         <oasis:entry colname="col4">1.22</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">23.52</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.92</oasis:entry>
         <oasis:entry colname="col7">0.4</oasis:entry>
         <oasis:entry colname="col8">0.55</oasis:entry>
         <oasis:entry colname="col9">0.76</oasis:entry>
         <oasis:entry colname="col10">0.82</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Baseline Reanalysis</oasis:entry>
         <oasis:entry colname="col2">0.69</oasis:entry>
         <oasis:entry colname="col3">0.39</oasis:entry>
         <oasis:entry colname="col4">4.36</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M168" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">85.12</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.93</oasis:entry>
         <oasis:entry colname="col7">0.41</oasis:entry>
         <oasis:entry colname="col8">0.58</oasis:entry>
         <oasis:entry colname="col9">0.77</oasis:entry>
         <oasis:entry colname="col10">0.82</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Baseline Reanalysis (Complex Embedding)</oasis:entry>
         <oasis:entry colname="col2">0.69</oasis:entry>
         <oasis:entry colname="col3">0.51</oasis:entry>
         <oasis:entry colname="col4">2.26</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">44.95</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.93</oasis:entry>
         <oasis:entry colname="col7">0.36</oasis:entry>
         <oasis:entry colname="col8">0.57</oasis:entry>
         <oasis:entry colname="col9">0.77</oasis:entry>
         <oasis:entry colname="col10">0.82</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1" align="left">Sequential Forecast LSTM with <inline-formula><mml:math id="M170" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> (ComplexEmbedding)</oasis:entry>
         <oasis:entry colname="col2">0.70</oasis:entry>
         <oasis:entry colname="col3">0.63</oasis:entry>
         <oasis:entry colname="col4">0.73</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">13.36</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.94</oasis:entry>
         <oasis:entry colname="col7">0.42</oasis:entry>
         <oasis:entry colname="col8">0.59</oasis:entry>
         <oasis:entry colname="col9">0.81</oasis:entry>
         <oasis:entry colname="col10">0.86</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1" align="left">Sequential Forecast LSTM with <inline-formula><mml:math id="M172" display="inline"><mml:mi>q</mml:mi></mml:math></inline-formula> (SimpleEmbedding)</oasis:entry>
         <oasis:entry colname="col2">0.71</oasis:entry>
         <oasis:entry colname="col3">0.4</oasis:entry>
         <oasis:entry colname="col4">4.27</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">78.3</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0.96</oasis:entry>
         <oasis:entry colname="col7">0.42</oasis:entry>
         <oasis:entry colname="col8">0.58</oasis:entry>
         <oasis:entry colname="col9">0.81</oasis:entry>
         <oasis:entry colname="col10">0.88</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>


</app>

<app id="App1.Ch1.S7">
  <label>Appendix G</label><title>Computational Resources</title>
      <p id="d2e6435">All conducted experiments were trained on a NVIDIA RTX4090 graphics processing unit, with wall times varying between several minutes to approximately 1 h for one model run, depending on the size of the input vector in the model and the model architecture itself. Although it is common practice in hyperparameter optimisation to run the same settings three times with different seedings, we have only run the tuning with one seed at a time due to computational constraints.</p>
</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d2e6442">All experiments have been conducted with a forked version of the NeuralHydrology library (Kratzert et al., 2022), available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.17293199" ext-link-type="DOI">10.5281/zenodo.17293199</ext-link> (Konold, 2025). The Extended LamaH-CE dataset is available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.17119635" ext-link-type="DOI">10.5281/zenodo.17119635</ext-link> (Konold et al., 2025b). The code to create the analysis and figures is available at <uri>https://github.com/conestone/biascast</uri> (last access: 30 July 2026). All trained models with its configuration files and saved weights are available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.17241922" ext-link-type="DOI">10.5281/zenodo.17241922</ext-link> (Konold et al., 2025a).</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e6460">OK and MF developed the idea and conceptualized the paper with support of KS. OK designed the experiments, build the model pipelines and did the LSTM simulations. The results were analyzed and interpreted by OK, MF and KS. PP contributed by writing the code of the hyperparameter optimization. CK provided the code and knowledge to extend the LamaH-CE dataset. The manuscript was written by OK and reviewed and corrected by MF and KS.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e6466">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e6472">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e6478">The need for appropriate LSTM training methods in a dedicated flood forecasting setting emerged in the KiHoRiMo2 project funded by the Federal Austrian Ministry (BMLUK). The paper was supported by the HADRIAN DocSchool from BOKU University, Vienna.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e6483">This paper was edited by Micha Werner and reviewed by Gwyneth Matthews and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bib1"><label>1</label><mixed-citation>Abbas, A., Yang, Y., Pan, M., Tramblay, Y., Shen, C., Ji, H., Gebrechorkos, S. H., Pappenberger, F., Pyo, J., Feng, D., Huffman, G., Nguyen, P., Massari, C., Brocca, L., Tan, J., and Beck, H. E.: Comprehensive Global Assessment of 24 Gridded Precipitation Datasets Across 18 428 Catchments Using Hydrological Modeling, Hydrol. Earth Syst. Sci., 30, 3399–3423, <ext-link xlink:href="https://doi.org/10.5194/hess-30-3399-2026" ext-link-type="DOI">10.5194/hess-30-3399-2026</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bib2"><label>2</label><mixed-citation>Acuña Espinoza, E., Loritz, R., Kratzert, F., Klotz, D., Gauch, M., Álvarez Chaves, M., and Ehret, U.: Analyzing the generalization capabilities of a hybrid hydrological model for extrapolation to extreme events, Hydrol. Earth Syst. Sci., 29, 1277–1294, <ext-link xlink:href="https://doi.org/10.5194/hess-29-1277-2025" ext-link-type="DOI">10.5194/hess-29-1277-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib3"><label>3</label><mixed-citation>Ahmed, S. F., Alam, Md. S. B., Hassan, M., Rozbu, M. R., Ishtiak, T., Rafa, N., Mofijur, M., Shawkat Ali, A. B. M., and Gandomi, A. H.: Deep learning modelling techniques: current progress, applications, advantages, and challenges, Artif. Intell. Rev., 56, 13521–13617, <ext-link xlink:href="https://doi.org/10.1007/s10462-023-10466-8" ext-link-type="DOI">10.1007/s10462-023-10466-8</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bib4"><label>4</label><mixed-citation>Alfieri, L., Burek, P., Dutra, E., Krzeminski, B., Muraro, D., Thielen, J., and Pappenberger, F.: GloFAS – global ensemble streamflow forecasting and flood early warning, Hydrol. Earth Syst. Sci., 17, 1161–1175, <ext-link xlink:href="https://doi.org/10.5194/hess-17-1161-2013" ext-link-type="DOI">10.5194/hess-17-1161-2013</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bib5"><label>5</label><mixed-citation>Bárdossy, A., Kilsby, C., Birkinshaw, S., Wang, N., and Anwar, F.: Is Precipitation Responsible for the Most Hydrological Model Uncertainty?, Front. Water, 4, 836554, <ext-link xlink:href="https://doi.org/10.3389/frwa.2022.836554" ext-link-type="DOI">10.3389/frwa.2022.836554</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib6"><label>6</label><mixed-citation>Beck, H. E., Vergopolan, N., Pan, M., Levizzani, V., van Dijk, A. I. J. M., Weedon, G. P., Brocca, L., Pappenberger, F., Huffman, G. J., and Wood, E. F.: Global-scale evaluation of 22 precipitation datasets using gauge observations and hydrological modeling, Hydrol. Earth Syst. Sci., 21, 6201–6217, <ext-link xlink:href="https://doi.org/10.5194/hess-21-6201-2017" ext-link-type="DOI">10.5194/hess-21-6201-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bib7"><label>7</label><mixed-citation>Beck, H. E., Wood, E. F., Pan, M., Fisher, C. K., Miralles, D. G., Van Dijk, A. I. J. M., McVicar, T. R., and Adler, R. F.: MSWEP V2 Global 3-Hourly 0.1° Precipitation: Methodology and Quantitative Assessment, B. Am. Meteorol. Soc., 100, 473–500, <ext-link xlink:href="https://doi.org/10.1175/BAMS-D-17-0138.1" ext-link-type="DOI">10.1175/BAMS-D-17-0138.1</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib8"><label>8</label><mixed-citation>Beven, K.: Rainfall‐Runoff Modelling: The Primer, 1st edn., Wiley, <ext-link xlink:href="https://doi.org/10.1002/9781119951001" ext-link-type="DOI">10.1002/9781119951001</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib9"><label>9</label><mixed-citation>Chen, X., Zhang, L., Gippel, C. J., Shan, L., Chen, S., and Yang, W.: Uncertainty of Flood Forecasting Based on Radar Rainfall Data Assimilation, Adv. Meteorol., 2016, 1–12, <ext-link xlink:href="https://doi.org/10.1155/2016/2710457" ext-link-type="DOI">10.1155/2016/2710457</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib10"><label>10</label><mixed-citation>Cornes, R. C., Van Der Schrier, G., Van Den Besselaar, E. J. M., and Jones, P. D.: An Ensemble Version of the E‐OBS Temperature and Precipitation Data Sets, J. Geophys. Res.-Atmos., 123, 9391–9409, <ext-link xlink:href="https://doi.org/10.1029/2017JD028200" ext-link-type="DOI">10.1029/2017JD028200</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib11"><label>11</label><mixed-citation>De Oliveira, D. Y. and Vrugt, J. A.: The Treatment of Uncertainty in Hydrometric Observations: A Probabilistic Description of Streamflow Records, Water Resour. Res., 58, e2022WR032263, <ext-link xlink:href="https://doi.org/10.1029/2022WR032263" ext-link-type="DOI">10.1029/2022WR032263</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib12"><label>12</label><mixed-citation>ECMWF: European Center for Medium-Range Weather Forecast High Resoultion Forecast (ECMWF-HRES), ECMWF [data set], <uri>https://www.ecmwf.int/en/forecasts/dataset/operational-archive</uri> (last access: 5 October 2025), 2025.</mixed-citation></ref>
      <ref id="bib1.bib13"><label>13</label><mixed-citation>Frazier, P. I.: A Tutorial on Bayesian Optimization, ARXIV [preprint], <ext-link xlink:href="https://doi.org/10.48550/ARXIV.1807.02811" ext-link-type="DOI">10.48550/ARXIV.1807.02811</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib14"><label>14</label><mixed-citation>Gauch, M., Kratzert, F., Klotz, D., Nearing, G., Cohen, D., and Gilon, O.: How to deal w___ missing input data, Hydrol. Earth Syst. Sci., 29, 6221–6235, <ext-link xlink:href="https://doi.org/10.5194/hess-29-6221-2025" ext-link-type="DOI">10.5194/hess-29-6221-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib15"><label>15</label><mixed-citation>Gers, F. A., Schmidhuber, J., and Cummins, F.: Learning to Forget: Continual Prediction with LSTM, Neural Comput., 12, 2451–2471, <ext-link xlink:href="https://doi.org/10.1162/089976600300015015" ext-link-type="DOI">10.1162/089976600300015015</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bib16"><label>16</label><mixed-citation>Goodfellow, I., Bengio, Y., and Courville, A.: Deep learning, The MIT press, Cambridge, Mass, <uri>https://www.deeplearningbook.org/</uri> (last access: 24 June 2025), 2016.</mixed-citation></ref>
      <ref id="bib1.bib17"><label>17</label><mixed-citation>Guo, K., Guan, M., and Yu, D.: Urban surface water flood modelling – a comprehensive review of current models and future challenges, Hydrol. Earth Syst. Sci., 25, 2843–2860, <ext-link xlink:href="https://doi.org/10.5194/hess-25-2843-2021" ext-link-type="DOI">10.5194/hess-25-2843-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib18"><label>18</label><mixed-citation>Haiden, T., Kann, A., Wittmann, C., Pistotnik, G., Bica, B., and Gruber, C.: The Integrated Nowcasting through Comprehensive Analysis (INCA) System and Its Validation over the Eastern Alpine Region, Weather Forecast., 26, 166–183, <ext-link xlink:href="https://doi.org/10.1175/2010WAF2222451.1" ext-link-type="DOI">10.1175/2010WAF2222451.1</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib19"><label>19</label><mixed-citation>Haiden, T., Janousek, M., Vitart, F., Tanguy, M., Prates, F., and Chevalier, M.: Evaluation of ECMWF forecasts, ECMWF, technical memorandum, <ext-link xlink:href="https://doi.org/10.21957/52F2F31351" ext-link-type="DOI">10.21957/52F2F31351</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib20"><label>20</label><mixed-citation>Han, L., Chen, M., Chen, K., Chen, H., Zhang, Y., Lu, B., Song, L., and Qin, R.: A Deep Learning Method for Bias Correction of ECMWF 24–240 h Forecasts, Adv. Atmos. Sci., 38, 1444–1459, <ext-link xlink:href="https://doi.org/10.1007/s00376-021-0215-y" ext-link-type="DOI">10.1007/s00376-021-0215-y</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib21"><label>21</label><mixed-citation>Herrnegger, M., Nachtnebel, H. P., and Schulz, K.: From runoff to rainfall: inverse rainfall–runoff modelling in a high temporal resolution, Hydrol. Earth Syst. Sci., 19, 4619–4639, <ext-link xlink:href="https://doi.org/10.5194/hess-19-4619-2015" ext-link-type="DOI">10.5194/hess-19-4619-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bib22"><label>22</label><mixed-citation>Hess, R.: Statistical postprocessing of ensemble forecasts for severe weather at Deutscher Wetterdienst, Nonlin. Processes Geophys., 27, 473–487, <ext-link xlink:href="https://doi.org/10.5194/npg-27-473-2020" ext-link-type="DOI">10.5194/npg-27-473-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib23"><label>23</label><mixed-citation>Hochreiter, S. and Schmidhuber, J.: Long Short-Term Memory, Neural Comput., 9, 1735–1780, <ext-link xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735" ext-link-type="DOI">10.1162/neco.1997.9.8.1735</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bib24"><label>24</label><mixed-citation>Hosna, A., Merry, E., Gyalmo, J., Alom, Z., Aung, Z., and Azim, M. A.: Transfer learning: a friendly introduction, J. Big Data, 9, 102, <ext-link xlink:href="https://doi.org/10.1186/s40537-022-00652-w" ext-link-type="DOI">10.1186/s40537-022-00652-w</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib25"><label>25</label><mixed-citation>Hunt, K. M. R., Matthews, G. R., Pappenberger, F., and Prudhomme, C.: Using a long short-term memory (LSTM) neural network to boost river streamflow forecasts over the western United States, Hydrol. Earth Syst. Sci., 26, 5449–5472, <ext-link xlink:href="https://doi.org/10.5194/hess-26-5449-2022" ext-link-type="DOI">10.5194/hess-26-5449-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib26"><label>26</label><mixed-citation>Irani, H., Ghahremani, Y., Kermani, A., and Metsis, V.: Time Series Embedding Methods for Classification Tasks: A Review, ARXIV [preprint], <ext-link xlink:href="https://doi.org/10.48550/ARXIV.2501.13392" ext-link-type="DOI">10.48550/ARXIV.2501.13392</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib27"><label>27</label><mixed-citation>Kirchner, J. W.: Catchments as simple dynamical systems: Catchment characterization, rainfall‐runoff modeling, and doing hydrology backward, Water Resour. Res., 45, 2008WR006912, <ext-link xlink:href="https://doi.org/10.1029/2008WR006912" ext-link-type="DOI">10.1029/2008WR006912</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bib28"><label>28</label><mixed-citation>Klingler, C., Schulz, K., and Herrnegger, M.: LamaH-CE: LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe, Earth Syst. Sci. Data, 13, 4529–4565, <ext-link xlink:href="https://doi.org/10.5194/essd-13-4529-2021" ext-link-type="DOI">10.5194/essd-13-4529-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib29"><label>29</label><mixed-citation>Ko, C.-M., Jeong, Y. Y., Lee, Y.-M., and Kim, B.-S.: The Development of a Quantitative Precipitation Forecast Correction Technique Based on Machine Learning for Hydrological Applications, Atmosphere, 11, 111, <ext-link xlink:href="https://doi.org/10.3390/atmos11010111" ext-link-type="DOI">10.3390/atmos11010111</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib30"><label>30</label><mixed-citation>Konold, O.: conestone/biascast: v1.0 (Version v1.0), Zenodo [software], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17293199" ext-link-type="DOI">10.5281/zenodo.17293199</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib31"><label>31</label><mixed-citation>Konold, O., Feigl, M., and Schulz, K.: Experimental Setups and Results for “BiasCast: Learning and adjusting real time biases from meteorological forecasts to enhance runoff predictions”, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17241922" ext-link-type="DOI">10.5281/zenodo.17241922</ext-link>, 2025a.</mixed-citation></ref>
      <ref id="bib1.bib32"><label>32</label><mixed-citation>Konold, O., Klingler, C., Feigl, M., Herrnegger, M., and Schulz, K.: Extended LamaH-CE: LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe (Version 1.0), Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.17119635" ext-link-type="DOI">10.5281/zenodo.17119635</ext-link>, 2025b.</mixed-citation></ref>
      <ref id="bib1.bib33"><label>33</label><mixed-citation>Kratzert, F., Klotz, D., Brenner, C., Schulz, K., and Herrnegger, M.: Rainfall–runoff modelling using Long Short-Term Memory (LSTM) networks, Hydrol. Earth Syst. Sci., 22, 6005–6022, <ext-link xlink:href="https://doi.org/10.5194/hess-22-6005-2018" ext-link-type="DOI">10.5194/hess-22-6005-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib34"><label>34</label><mixed-citation>Kratzert, F., Herrnegger, M., Klotz, D., Hochreiter, S., and Klambauer, G.: NeuralHydrology – Interpreting LSTMs in Hydrology, in: Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, vol. 11700, edited by: Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K., and Müller, K.-R., Springer International Publishing, Cham, 347–362, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-28954-6_19" ext-link-type="DOI">10.1007/978-3-030-28954-6_19</ext-link>, 2019a.</mixed-citation></ref>
      <ref id="bib1.bib35"><label>35</label><mixed-citation>Kratzert, F., Klotz, D., Herrnegger, M., Sampson, A. K., Hochreiter, S., and Nearing, G. S.: Toward Improved Predictions in Ungauged Basins: Exploiting the Power of Machine Learning, Water Resour. Res., 55, 11344–11354, <ext-link xlink:href="https://doi.org/10.1029/2019WR026065" ext-link-type="DOI">10.1029/2019WR026065</ext-link>, 2019b.</mixed-citation></ref>
      <ref id="bib1.bib36"><label>36</label><mixed-citation>Kratzert, F., Klotz, D., Shalev, G., Klambauer, G., Hochreiter, S., and Nearing, G.: Towards learning universal, regional, and local hydrological behaviors via machine learning applied to large-sample datasets, Hydrol. Earth Syst. Sci., 23, 5089–5110, <ext-link xlink:href="https://doi.org/10.5194/hess-23-5089-2019" ext-link-type="DOI">10.5194/hess-23-5089-2019</ext-link>, 2019c.</mixed-citation></ref>
      <ref id="bib1.bib37"><label>37</label><mixed-citation>Kratzert, F., Klotz, D., Hochreiter, S., and Nearing, G. S.: A note on leveraging synergy in multiple meteorological data sets with deep learning for rainfall–runoff modeling, Hydrol. Earth Syst. Sci., 25, 2685–2703, <ext-link xlink:href="https://doi.org/10.5194/hess-25-2685-2021" ext-link-type="DOI">10.5194/hess-25-2685-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib38"><label>38</label><mixed-citation>Kratzert, F., Gauch, M., Nearing, G., and Klotz, D.: NeuralHydrology – A Python library for Deep Learningresearch in hydrology, JOSS, 7, 4050, <ext-link xlink:href="https://doi.org/10.21105/joss.04050" ext-link-type="DOI">10.21105/joss.04050</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib39"><label>39</label><mixed-citation>Lavers, D. A., Harrigan, S., and Prudhomme, C.: Precipitation Biases in the ECMWF Integrated Forecasting System, J. Hydrometeorol., 22, 1187–1198, <ext-link xlink:href="https://doi.org/10.1175/JHM-D-20-0308.1" ext-link-type="DOI">10.1175/JHM-D-20-0308.1</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib40"><label>40</label><mixed-citation>Lenderink, G., Buishand, A., and van Deursen, W.: Estimates of future discharges of the river Rhine using two scenario methodologies: direct versus delta approach, Hydrol. Earth Syst. Sci., 11, 1145–1159, <ext-link xlink:href="https://doi.org/10.5194/hess-11-1145-2007" ext-link-type="DOI">10.5194/hess-11-1145-2007</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bib41"><label>41</label><mixed-citation>Mao, R., Wang, L., Zhou, J., Li, X., Qi, J., and Zhang, X.: Evaluation of Various Precipitation Products Using Ground-Based Discharge Observation at the Nujiang River Basin, China, Water, 11, 2308, <ext-link xlink:href="https://doi.org/10.3390/w11112308" ext-link-type="DOI">10.3390/w11112308</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bib42"><label>42</label><mixed-citation>Miralles, D. G., Holmes, T. R. H., De Jeu, R. A. M., Gash, J. H., Meesters, A. G. C. A., and Dolman, A. J.: Global land-surface evaporation estimated from satellite-based observations, Hydrol. Earth Syst. Sci., 15, 453–469, <ext-link xlink:href="https://doi.org/10.5194/hess-15-453-2011" ext-link-type="DOI">10.5194/hess-15-453-2011</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bib43"><label>43</label><mixed-citation>Muñoz-Sabater, J., Dutra, E., Agustí-Panareda, A., Albergel, C., Arduini, G., Balsamo, G., Boussetta, S., Choulga, M., Harrigan, S., Hersbach, H., Martens, B., Miralles, D. G., Piles, M., Rodríguez-Fernández, N. J., Zsoter, E., Buontempo, C., and Thépaut, J.-N.: ERA5-Land: a state-of-the-art global reanalysis dataset for land applications, Earth Syst. Sci. Data, 13, 4349–4383, <ext-link xlink:href="https://doi.org/10.5194/essd-13-4349-2021" ext-link-type="DOI">10.5194/essd-13-4349-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib44"><label>44</label><mixed-citation>Nash, J. E. and Sutcliffe, J. V.: River flow forecasting through conceptual models part I – A discussion of principles, J. Hydrol., 10, 282–290, <ext-link xlink:href="https://doi.org/10.1016/0022-1694(70)90255-6" ext-link-type="DOI">10.1016/0022-1694(70)90255-6</ext-link>, 1970.</mixed-citation></ref>
      <ref id="bib1.bib45"><label>45</label><mixed-citation>Nearing, G., Cohen, D., Dube, V., Gauch, M., Gilon, O., Harrigan, S., Hassidim, A., Klotz, D., Kratzert, F., Metzger, A., Nevo, S., Pappenberger, F., Prudhomme, C., Shalev, G., Shenzis, S., Tekalign, T. Y., Weitzner, D., and Matias, Y.: Global prediction of extreme floods in ungauged watersheds, Nature, 627, 559–563, <ext-link xlink:href="https://doi.org/10.1038/s41586-024-07145-1" ext-link-type="DOI">10.1038/s41586-024-07145-1</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bib46"><label>46</label><mixed-citation>Nearing, G. S., Kratzert, F., Sampson, A. K., Pelissier, C. S., Klotz, D., Frame, J. M., Prieto, C., and Gupta, H. V.: What Role Does Hydrological Science Play in the Age of Machine Learning?, Water Resour. Res., 57, e2020WR028091, <ext-link xlink:href="https://doi.org/10.1029/2020WR028091" ext-link-type="DOI">10.1029/2020WR028091</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bib47"><label>47</label><mixed-citation>Nearing, G. S., Klotz, D., Frame, J. M., Gauch, M., Gilon, O., Kratzert, F., Sampson, A. K., Shalev, G., and Nevo, S.: Technical note: Data assimilation and autoregression for using near-real-time streamflow observations in long short-term memory networks, Hydrol. Earth Syst. Sci., 26, 5493–5513, <ext-link xlink:href="https://doi.org/10.5194/hess-26-5493-2022" ext-link-type="DOI">10.5194/hess-26-5493-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bib48"><label>48</label><mixed-citation>Nester, T., Komma, J., Viglione, A., and Blöschl, G.: Flood forecast errors and ensemble spread – A case study, Water Resour. Res., 48, 2011WR011649, <ext-link xlink:href="https://doi.org/10.1029/2011WR011649" ext-link-type="DOI">10.1029/2011WR011649</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib49"><label>49</label><mixed-citation>Newman, A. J., Clark, M. P., Sampson, K., Wood, A., Hay, L. E., Bock, A., Viger, R. J., Blodgett, D., Brekke, L., Arnold, J. R., Hopson, T., and Duan, Q.: Development of a large-sample watershed-scale hydrometeorological data set for the contiguous USA: data set characteristics and assessment of regional variability in hydrologic model performance, Hydrol. Earth Syst. Sci., 19, 209–223, <ext-link xlink:href="https://doi.org/10.5194/hess-19-209-2015" ext-link-type="DOI">10.5194/hess-19-209-2015</ext-link>, 2015. </mixed-citation></ref>
      <ref id="bib1.bib50"><label>50</label><mixed-citation>Pan, S. J. and Yang, Q.: A Survey on Transfer Learning, IEEE Trans. Knowl. Data Eng., 22, 1345–1359, <ext-link xlink:href="https://doi.org/10.1109/tkde.2009.191" ext-link-type="DOI">10.1109/tkde.2009.191</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bib51"><label>51</label><mixed-citation>Seibert, J., Vis, M. J. P., Lewis, E., and Van Meerveld, H. J.: Upper and lower benchmarks in hydrological modelling, Hydrol. Process., 32, 1120–1125, <ext-link xlink:href="https://doi.org/10.1002/hyp.11476" ext-link-type="DOI">10.1002/hyp.11476</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bib52"><label>52</label><mixed-citation>Snoek, J., Larochelle, H., and Adams, R. P.: Practical Bayesian Optimization of Machine Learning Algorithms, ARXIV [preprint], <ext-link xlink:href="https://doi.org/10.48550/ARXIV.1206.2944" ext-link-type="DOI">10.48550/ARXIV.1206.2944</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bib53"><label>53</label><mixed-citation>Tran, C. K., Dang, N. D., Nguyen, D. M., Nguyen, B. T. N., Le, B. T. H., Vo, H. C., and La, H. P.: Real-time flood forecasting using time-varying parameter hydrological model: case study for Ta Trach reservoir, Appl. Water Sci., 15, 152, <ext-link xlink:href="https://doi.org/10.1007/s13201-025-02503-4" ext-link-type="DOI">10.1007/s13201-025-02503-4</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bib54"><label>54</label><mixed-citation>Villani, C.: The Wasserstein distances, in: Optimal Transport, vol. 338, Springer Berlin Heidelberg, Berlin, Heidelberg, 93–111, <ext-link xlink:href="https://doi.org/10.1007/978-3-540-71050-9_6" ext-link-type="DOI">10.1007/978-3-540-71050-9_6</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bib55"><label>55</label><mixed-citation>Villarini, G., Mandapaka, P. V., Krajewski, W. F., and Moore, R. J.: Rainfall and sampling uncertainties: A rain gauge perspective, J. Geophys. Res., 113, 2007JD009214, <ext-link xlink:href="https://doi.org/10.1029/2007JD009214" ext-link-type="DOI">10.1029/2007JD009214</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bib56"><label>56</label><mixed-citation>Weiss, K., Khoshgoftaar, T. M., and Wang, D.: A survey of transfer learning, J. Big Data, 3, <ext-link xlink:href="https://doi.org/10.1186/s40537-016-0043-6" ext-link-type="DOI">10.1186/s40537-016-0043-6</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bib57"><label>57</label><mixed-citation>Yang, D., Ishida, S., Goodison, B. E., and Gunther, T.: Bias correction of daily precipitation measurements for Greenland, J. Geophys. Res., 104, 6171–6181, <ext-link xlink:href="https://doi.org/10.1029/1998JD200110" ext-link-type="DOI">10.1029/1998JD200110</ext-link>, 1999.</mixed-citation></ref>
      <ref id="bib1.bib58"><label>58</label><mixed-citation>Zhang, C., Zeng, J., Wang, H., Ma, L., and Chu, H.: Correction model for rainfall forecasts using the LSTM with multiple meteorological factors, Meteorol. Appl., 27, e1852, <ext-link xlink:href="https://doi.org/10.1002/met.1852" ext-link-type="DOI">10.1002/met.1852</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bib59"><label>59</label><mixed-citation>Zhuang, F., Qi, Z., Duan, K., Xi, D., Zhu, Y., Zhu, H., Xiong, H., and He, Q.: A Comprehensive Survey on Transfer Learning, Proc. IEEE, 109, 43–76, <ext-link xlink:href="https://doi.org/10.1109/JPROC.2020.3004555" ext-link-type="DOI">10.1109/JPROC.2020.3004555</ext-link>, 2021.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>BiasCast: learning and adjusting real time biases from meteorological forecasts to enhance runoff predictions</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>1</label><mixed-citation>
      
Abbas, A., Yang, Y., Pan, M., Tramblay, Y., Shen, C., Ji, H., Gebrechorkos, S. H., Pappenberger, F., Pyo, J., Feng, D., Huffman, G., Nguyen, P., Massari, C., Brocca, L., Tan, J., and Beck, H. E.: Comprehensive Global Assessment of 24 Gridded Precipitation Datasets Across 18 428 Catchments Using Hydrological Modeling, Hydrol. Earth Syst. Sci., 30, 3399–3423, <a href="https://doi.org/10.5194/hess-30-3399-2026" target="_blank">https://doi.org/10.5194/hess-30-3399-2026</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>2</label><mixed-citation>
      
Acuña Espinoza, E., Loritz, R., Kratzert, F., Klotz, D., Gauch, M., Álvarez Chaves, M., and Ehret, U.: Analyzing the generalization capabilities of a hybrid hydrological model for extrapolation to extreme events, Hydrol. Earth Syst. Sci., 29, 1277–1294, <a href="https://doi.org/10.5194/hess-29-1277-2025" target="_blank">https://doi.org/10.5194/hess-29-1277-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>3</label><mixed-citation>
      
Ahmed, S. F., Alam, Md. S. B., Hassan, M., Rozbu, M. R., Ishtiak, T., Rafa, N., Mofijur, M., Shawkat Ali, A. B. M., and Gandomi, A. H.: Deep learning modelling techniques: current progress, applications, advantages, and challenges, Artif. Intell. Rev., 56, 13521–13617, <a href="https://doi.org/10.1007/s10462-023-10466-8" target="_blank">https://doi.org/10.1007/s10462-023-10466-8</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>4</label><mixed-citation>
      
Alfieri, L., Burek, P., Dutra, E., Krzeminski, B., Muraro, D., Thielen, J., and Pappenberger, F.: GloFAS – global ensemble streamflow forecasting and flood early warning, Hydrol. Earth Syst. Sci., 17, 1161–1175, <a href="https://doi.org/10.5194/hess-17-1161-2013" target="_blank">https://doi.org/10.5194/hess-17-1161-2013</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>5</label><mixed-citation>
      
Bárdossy, A., Kilsby, C., Birkinshaw, S., Wang, N., and Anwar, F.: Is Precipitation Responsible for the Most Hydrological Model Uncertainty?, Front. Water, 4, 836554, <a href="https://doi.org/10.3389/frwa.2022.836554" target="_blank">https://doi.org/10.3389/frwa.2022.836554</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>6</label><mixed-citation>
      
Beck, H. E., Vergopolan, N., Pan, M., Levizzani, V., van Dijk, A. I. J. M., Weedon, G. P., Brocca, L., Pappenberger, F., Huffman, G. J., and Wood, E. F.: Global-scale evaluation of 22 precipitation datasets using gauge observations and hydrological modeling, Hydrol. Earth Syst. Sci., 21, 6201–6217, <a href="https://doi.org/10.5194/hess-21-6201-2017" target="_blank">https://doi.org/10.5194/hess-21-6201-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>7</label><mixed-citation>
      
Beck, H. E., Wood, E. F., Pan, M., Fisher, C. K., Miralles, D. G., Van Dijk, A. I. J. M., McVicar, T. R., and Adler, R. F.: MSWEP V2 Global 3-Hourly 0.1° Precipitation: Methodology and Quantitative Assessment, B. Am. Meteorol. Soc., 100, 473–500, <a href="https://doi.org/10.1175/BAMS-D-17-0138.1" target="_blank">https://doi.org/10.1175/BAMS-D-17-0138.1</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>8</label><mixed-citation>
      
Beven, K.: Rainfall‐Runoff Modelling: The Primer, 1st edn., Wiley, <a href="https://doi.org/10.1002/9781119951001" target="_blank">https://doi.org/10.1002/9781119951001</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>9</label><mixed-citation>
      
Chen, X., Zhang, L., Gippel, C. J., Shan, L., Chen, S., and Yang, W.: Uncertainty of Flood Forecasting Based on Radar Rainfall Data Assimilation, Adv. Meteorol., 2016, 1–12, <a href="https://doi.org/10.1155/2016/2710457" target="_blank">https://doi.org/10.1155/2016/2710457</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>10</label><mixed-citation>
      
Cornes, R. C., Van Der Schrier, G., Van Den Besselaar, E. J. M., and Jones, P. D.: An Ensemble Version of the E‐OBS Temperature and Precipitation Data Sets, J. Geophys. Res.-Atmos., 123, 9391–9409, <a href="https://doi.org/10.1029/2017JD028200" target="_blank">https://doi.org/10.1029/2017JD028200</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>11</label><mixed-citation>
      
De Oliveira, D. Y. and Vrugt, J. A.: The Treatment of Uncertainty in Hydrometric Observations: A Probabilistic Description of Streamflow Records, Water Resour. Res., 58, e2022WR032263, <a href="https://doi.org/10.1029/2022WR032263" target="_blank">https://doi.org/10.1029/2022WR032263</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>12</label><mixed-citation>
      
ECMWF: European Center for Medium-Range Weather Forecast High Resoultion Forecast (ECMWF-HRES), ECMWF [data set], <a href="https://www.ecmwf.int/en/forecasts/dataset/operational-archive" target="_blank"/> (last access: 5 October 2025), 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>13</label><mixed-citation>
      
Frazier, P. I.: A Tutorial on Bayesian Optimization, ARXIV [preprint], <a href="https://doi.org/10.48550/ARXIV.1807.02811" target="_blank">https://doi.org/10.48550/ARXIV.1807.02811</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>14</label><mixed-citation>
      
Gauch, M., Kratzert, F., Klotz, D., Nearing, G., Cohen, D., and Gilon, O.: How to deal w___ missing input data, Hydrol. Earth Syst. Sci., 29, 6221–6235, <a href="https://doi.org/10.5194/hess-29-6221-2025" target="_blank">https://doi.org/10.5194/hess-29-6221-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>15</label><mixed-citation>
      
Gers, F. A., Schmidhuber, J., and Cummins, F.: Learning to Forget: Continual Prediction with LSTM, Neural Comput., 12, 2451–2471, <a href="https://doi.org/10.1162/089976600300015015" target="_blank">https://doi.org/10.1162/089976600300015015</a>, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>16</label><mixed-citation>
      
Goodfellow, I., Bengio, Y., and Courville, A.: Deep learning, The MIT press, Cambridge, Mass, <a href="https://www.deeplearningbook.org/" target="_blank"/> (last access: 24 June 2025), 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>17</label><mixed-citation>
      
Guo, K., Guan, M., and Yu, D.: Urban surface water flood modelling – a comprehensive review of current models and future challenges, Hydrol. Earth Syst. Sci., 25, 2843–2860, <a href="https://doi.org/10.5194/hess-25-2843-2021" target="_blank">https://doi.org/10.5194/hess-25-2843-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>18</label><mixed-citation>
      
Haiden, T., Kann, A., Wittmann, C., Pistotnik, G., Bica, B., and Gruber, C.: The Integrated Nowcasting through Comprehensive Analysis (INCA) System and Its Validation over the Eastern Alpine Region, Weather Forecast., 26, 166–183, <a href="https://doi.org/10.1175/2010WAF2222451.1" target="_blank">https://doi.org/10.1175/2010WAF2222451.1</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>19</label><mixed-citation>
      
Haiden, T., Janousek, M., Vitart, F., Tanguy, M., Prates, F., and Chevalier, M.: Evaluation of ECMWF forecasts, ECMWF, technical memorandum, <a href="https://doi.org/10.21957/52F2F31351" target="_blank">https://doi.org/10.21957/52F2F31351</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>20</label><mixed-citation>
      
Han, L., Chen, M., Chen, K., Chen, H., Zhang, Y., Lu, B., Song, L., and Qin, R.: A Deep Learning Method for Bias Correction of ECMWF 24–240&thinsp;h Forecasts, Adv. Atmos. Sci., 38, 1444–1459, <a href="https://doi.org/10.1007/s00376-021-0215-y" target="_blank">https://doi.org/10.1007/s00376-021-0215-y</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>21</label><mixed-citation>
      
Herrnegger, M., Nachtnebel, H. P., and Schulz, K.: From runoff to rainfall: inverse rainfall–runoff modelling in a high temporal resolution, Hydrol. Earth Syst. Sci., 19, 4619–4639, <a href="https://doi.org/10.5194/hess-19-4619-2015" target="_blank">https://doi.org/10.5194/hess-19-4619-2015</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>22</label><mixed-citation>
      
Hess, R.: Statistical postprocessing of ensemble forecasts for severe weather at Deutscher Wetterdienst, Nonlin. Processes Geophys., 27, 473–487, <a href="https://doi.org/10.5194/npg-27-473-2020" target="_blank">https://doi.org/10.5194/npg-27-473-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>23</label><mixed-citation>
      
Hochreiter, S. and Schmidhuber, J.: Long Short-Term Memory, Neural Comput., 9, 1735–1780, <a href="https://doi.org/10.1162/neco.1997.9.8.1735" target="_blank">https://doi.org/10.1162/neco.1997.9.8.1735</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>24</label><mixed-citation>
      
Hosna, A., Merry, E., Gyalmo, J., Alom, Z., Aung, Z., and Azim, M. A.: Transfer learning: a friendly introduction, J. Big Data, 9, 102, <a href="https://doi.org/10.1186/s40537-022-00652-w" target="_blank">https://doi.org/10.1186/s40537-022-00652-w</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>25</label><mixed-citation>
      
Hunt, K. M. R., Matthews, G. R., Pappenberger, F., and Prudhomme, C.: Using a long short-term memory (LSTM) neural network to boost river streamflow forecasts over the western United States, Hydrol. Earth Syst. Sci., 26, 5449–5472, <a href="https://doi.org/10.5194/hess-26-5449-2022" target="_blank">https://doi.org/10.5194/hess-26-5449-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>26</label><mixed-citation>
      
Irani, H., Ghahremani, Y., Kermani, A., and Metsis, V.: Time Series Embedding Methods for Classification Tasks: A Review, ARXIV [preprint], <a href="https://doi.org/10.48550/ARXIV.2501.13392" target="_blank">https://doi.org/10.48550/ARXIV.2501.13392</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>27</label><mixed-citation>
      
Kirchner, J. W.: Catchments as simple dynamical systems: Catchment characterization, rainfall‐runoff modeling, and doing hydrology backward, Water Resour. Res., 45, 2008WR006912, <a href="https://doi.org/10.1029/2008WR006912" target="_blank">https://doi.org/10.1029/2008WR006912</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>28</label><mixed-citation>
      
Klingler, C., Schulz, K., and Herrnegger, M.: LamaH-CE: LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe, Earth Syst. Sci. Data, 13, 4529–4565, <a href="https://doi.org/10.5194/essd-13-4529-2021" target="_blank">https://doi.org/10.5194/essd-13-4529-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>29</label><mixed-citation>
      
Ko, C.-M., Jeong, Y. Y., Lee, Y.-M., and Kim, B.-S.: The Development of a Quantitative Precipitation Forecast Correction Technique Based on Machine Learning for Hydrological Applications, Atmosphere, 11, 111, <a href="https://doi.org/10.3390/atmos11010111" target="_blank">https://doi.org/10.3390/atmos11010111</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>30</label><mixed-citation>
      
Konold, O.: conestone/biascast: v1.0 (Version v1.0), Zenodo [software], <a href="https://doi.org/10.5281/zenodo.17293199" target="_blank">https://doi.org/10.5281/zenodo.17293199</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>31</label><mixed-citation>
      
Konold, O., Feigl, M., and Schulz, K.: Experimental Setups and Results for “BiasCast: Learning and adjusting real time biases from meteorological forecasts to enhance runoff predictions”, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.17241922" target="_blank">https://doi.org/10.5281/zenodo.17241922</a>, 2025a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>32</label><mixed-citation>
      
Konold, O., Klingler, C., Feigl, M., Herrnegger, M., and Schulz, K.: Extended LamaH-CE: LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe (Version 1.0), Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.17119635" target="_blank">https://doi.org/10.5281/zenodo.17119635</a>, 2025b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>33</label><mixed-citation>
      
Kratzert, F., Klotz, D., Brenner, C., Schulz, K., and Herrnegger, M.: Rainfall–runoff modelling using Long Short-Term Memory (LSTM) networks, Hydrol. Earth Syst. Sci., 22, 6005–6022, <a href="https://doi.org/10.5194/hess-22-6005-2018" target="_blank">https://doi.org/10.5194/hess-22-6005-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>34</label><mixed-citation>
      
Kratzert, F., Herrnegger, M., Klotz, D., Hochreiter, S., and Klambauer, G.: NeuralHydrology – Interpreting LSTMs in Hydrology, in: Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, vol. 11700, edited by: Samek, W., Montavon, G., Vedaldi, A., Hansen, L. K., and Müller, K.-R., Springer International Publishing, Cham, 347–362, <a href="https://doi.org/10.1007/978-3-030-28954-6_19" target="_blank">https://doi.org/10.1007/978-3-030-28954-6_19</a>, 2019a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>35</label><mixed-citation>
      
Kratzert, F., Klotz, D., Herrnegger, M., Sampson, A. K., Hochreiter, S., and Nearing, G. S.: Toward Improved Predictions in Ungauged Basins: Exploiting the Power of Machine Learning, Water Resour. Res., 55, 11344–11354, <a href="https://doi.org/10.1029/2019WR026065" target="_blank">https://doi.org/10.1029/2019WR026065</a>, 2019b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>36</label><mixed-citation>
      
Kratzert, F., Klotz, D., Shalev, G., Klambauer, G., Hochreiter, S., and Nearing, G.: Towards learning universal, regional, and local hydrological behaviors via machine learning applied to large-sample datasets, Hydrol. Earth Syst. Sci., 23, 5089–5110, <a href="https://doi.org/10.5194/hess-23-5089-2019" target="_blank">https://doi.org/10.5194/hess-23-5089-2019</a>, 2019c.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>37</label><mixed-citation>
      
Kratzert, F., Klotz, D., Hochreiter, S., and Nearing, G. S.: A note on leveraging synergy in multiple meteorological data sets with deep learning for rainfall–runoff modeling, Hydrol. Earth Syst. Sci., 25, 2685–2703, <a href="https://doi.org/10.5194/hess-25-2685-2021" target="_blank">https://doi.org/10.5194/hess-25-2685-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>38</label><mixed-citation>
      
Kratzert, F., Gauch, M., Nearing, G., and Klotz, D.: NeuralHydrology – A Python library for Deep Learningresearch in hydrology, JOSS, 7, 4050, <a href="https://doi.org/10.21105/joss.04050" target="_blank">https://doi.org/10.21105/joss.04050</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>39</label><mixed-citation>
      
Lavers, D. A., Harrigan, S., and Prudhomme, C.: Precipitation Biases in the ECMWF Integrated Forecasting System, J. Hydrometeorol., 22, 1187–1198, <a href="https://doi.org/10.1175/JHM-D-20-0308.1" target="_blank">https://doi.org/10.1175/JHM-D-20-0308.1</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>40</label><mixed-citation>
      
Lenderink, G., Buishand, A., and van Deursen, W.: Estimates of future discharges of the river Rhine using two scenario methodologies: direct versus delta approach, Hydrol. Earth Syst. Sci., 11, 1145–1159, <a href="https://doi.org/10.5194/hess-11-1145-2007" target="_blank">https://doi.org/10.5194/hess-11-1145-2007</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>41</label><mixed-citation>
      
Mao, R., Wang, L., Zhou, J., Li, X., Qi, J., and Zhang, X.: Evaluation of Various Precipitation Products Using Ground-Based Discharge Observation at the Nujiang River Basin, China, Water, 11, 2308, <a href="https://doi.org/10.3390/w11112308" target="_blank">https://doi.org/10.3390/w11112308</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>42</label><mixed-citation>
      
Miralles, D. G., Holmes, T. R. H., De Jeu, R. A. M., Gash, J. H., Meesters, A. G. C. A., and Dolman, A. J.: Global land-surface evaporation estimated from satellite-based observations, Hydrol. Earth Syst. Sci., 15, 453–469, <a href="https://doi.org/10.5194/hess-15-453-2011" target="_blank">https://doi.org/10.5194/hess-15-453-2011</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>43</label><mixed-citation>
      
Muñoz-Sabater, J., Dutra, E., Agustí-Panareda, A., Albergel, C., Arduini, G., Balsamo, G., Boussetta, S., Choulga, M., Harrigan, S., Hersbach, H., Martens, B., Miralles, D. G., Piles, M., Rodríguez-Fernández, N. J., Zsoter, E., Buontempo, C., and Thépaut, J.-N.: ERA5-Land: a state-of-the-art global reanalysis dataset for land applications, Earth Syst. Sci. Data, 13, 4349–4383, <a href="https://doi.org/10.5194/essd-13-4349-2021" target="_blank">https://doi.org/10.5194/essd-13-4349-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>44</label><mixed-citation>
      
Nash, J. E. and Sutcliffe, J. V.: River flow forecasting through conceptual models part I – A discussion of principles, J. Hydrol., 10, 282–290, <a href="https://doi.org/10.1016/0022-1694(70)90255-6" target="_blank">https://doi.org/10.1016/0022-1694(70)90255-6</a>, 1970.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>45</label><mixed-citation>
      
Nearing, G., Cohen, D., Dube, V., Gauch, M., Gilon, O., Harrigan, S., Hassidim, A., Klotz, D., Kratzert, F., Metzger, A., Nevo, S., Pappenberger, F., Prudhomme, C., Shalev, G., Shenzis, S., Tekalign, T. Y., Weitzner, D., and Matias, Y.: Global prediction of extreme floods in ungauged watersheds, Nature, 627, 559–563, <a href="https://doi.org/10.1038/s41586-024-07145-1" target="_blank">https://doi.org/10.1038/s41586-024-07145-1</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>46</label><mixed-citation>
      
Nearing, G. S., Kratzert, F., Sampson, A. K., Pelissier, C. S., Klotz, D., Frame, J. M., Prieto, C., and Gupta, H. V.: What Role Does Hydrological Science Play in the Age of Machine Learning?, Water Resour. Res., 57, e2020WR028091, <a href="https://doi.org/10.1029/2020WR028091" target="_blank">https://doi.org/10.1029/2020WR028091</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>47</label><mixed-citation>
      
Nearing, G. S., Klotz, D., Frame, J. M., Gauch, M., Gilon, O., Kratzert, F., Sampson, A. K., Shalev, G., and Nevo, S.: Technical note: Data assimilation and autoregression for using near-real-time streamflow observations in long short-term memory networks, Hydrol. Earth Syst. Sci., 26, 5493–5513, <a href="https://doi.org/10.5194/hess-26-5493-2022" target="_blank">https://doi.org/10.5194/hess-26-5493-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>48</label><mixed-citation>
      
Nester, T., Komma, J., Viglione, A., and Blöschl, G.: Flood forecast errors and ensemble spread – A case study, Water Resour. Res., 48, 2011WR011649, <a href="https://doi.org/10.1029/2011WR011649" target="_blank">https://doi.org/10.1029/2011WR011649</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>49</label><mixed-citation>
      
Newman, A. J., Clark, M. P., Sampson, K., Wood, A., Hay, L. E., Bock, A., Viger, R. J., Blodgett, D., Brekke, L., Arnold, J. R., Hopson, T., and Duan, Q.: Development of a large-sample watershed-scale hydrometeorological data set for the contiguous USA: data set characteristics and assessment of regional variability in hydrologic model performance, Hydrol. Earth Syst. Sci., 19, 209–223, <a href="https://doi.org/10.5194/hess-19-209-2015" target="_blank">https://doi.org/10.5194/hess-19-209-2015</a>, 2015.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>50</label><mixed-citation>
      
Pan, S. J. and Yang, Q.: A Survey on Transfer Learning, IEEE Trans. Knowl. Data Eng., 22, 1345–1359, <a href="https://doi.org/10.1109/tkde.2009.191" target="_blank">https://doi.org/10.1109/tkde.2009.191</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>51</label><mixed-citation>
      
Seibert, J., Vis, M. J. P., Lewis, E., and Van Meerveld, H. J.: Upper and lower benchmarks in hydrological modelling, Hydrol. Process., 32, 1120–1125, <a href="https://doi.org/10.1002/hyp.11476" target="_blank">https://doi.org/10.1002/hyp.11476</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>52</label><mixed-citation>
      
Snoek, J., Larochelle, H., and Adams, R. P.: Practical Bayesian Optimization of Machine Learning Algorithms, ARXIV [preprint], <a href="https://doi.org/10.48550/ARXIV.1206.2944" target="_blank">https://doi.org/10.48550/ARXIV.1206.2944</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>53</label><mixed-citation>
      
Tran, C. K., Dang, N. D., Nguyen, D. M., Nguyen, B. T. N., Le, B. T. H., Vo, H. C., and La, H. P.: Real-time flood forecasting using time-varying parameter hydrological model: case study for Ta Trach reservoir, Appl. Water Sci., 15, 152, <a href="https://doi.org/10.1007/s13201-025-02503-4" target="_blank">https://doi.org/10.1007/s13201-025-02503-4</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>54</label><mixed-citation>
      
Villani, C.: The Wasserstein distances, in: Optimal Transport, vol. 338, Springer Berlin Heidelberg, Berlin, Heidelberg, 93–111, <a href="https://doi.org/10.1007/978-3-540-71050-9_6" target="_blank">https://doi.org/10.1007/978-3-540-71050-9_6</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>55</label><mixed-citation>
      
Villarini, G., Mandapaka, P. V., Krajewski, W. F., and Moore, R. J.: Rainfall and sampling uncertainties: A rain gauge perspective, J. Geophys. Res., 113, 2007JD009214, <a href="https://doi.org/10.1029/2007JD009214" target="_blank">https://doi.org/10.1029/2007JD009214</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>56</label><mixed-citation>
      
Weiss, K., Khoshgoftaar, T. M., and Wang, D.: A survey of transfer learning, J. Big Data, 3, <a href="https://doi.org/10.1186/s40537-016-0043-6" target="_blank">https://doi.org/10.1186/s40537-016-0043-6</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>57</label><mixed-citation>
      
Yang, D., Ishida, S., Goodison, B. E., and Gunther, T.: Bias correction of daily precipitation measurements for Greenland, J. Geophys. Res., 104, 6171–6181, <a href="https://doi.org/10.1029/1998JD200110" target="_blank">https://doi.org/10.1029/1998JD200110</a>, 1999.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>58</label><mixed-citation>
      
Zhang, C., Zeng, J., Wang, H., Ma, L., and Chu, H.: Correction model for rainfall forecasts using the LSTM with multiple meteorological factors, Meteorol. Appl., 27, e1852, <a href="https://doi.org/10.1002/met.1852" target="_blank">https://doi.org/10.1002/met.1852</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>59</label><mixed-citation>
      
Zhuang, F., Qi, Z., Duan, K., Xi, D., Zhu, Y., Zhu, H., Xiong, H., and He, Q.: A Comprehensive Survey on Transfer Learning, Proc. IEEE, 109, 43–76, <a href="https://doi.org/10.1109/JPROC.2020.3004555" target="_blank">https://doi.org/10.1109/JPROC.2020.3004555</a>, 2021.

    </mixed-citation></ref-html>--></article>
