<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">HESS</journal-id><journal-title-group>
    <journal-title>Hydrology and Earth System Sciences</journal-title>
    <abbrev-journal-title abbrev-type="publisher">HESS</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Hydrol. Earth Syst. Sci.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7938</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/hess-30-6207-2026</article-id><title-group><article-title>The ability of LSTM to model snowmelt versus rainfall generated floods</article-title><alt-title>The ability of LSTM to model snowmelt versus rainfall generated floods</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Bakke</surname><given-names>Sigrid Jørgensen</given-names></name>
          <email>sijb@nve.no</email>
        <ext-link>https://orcid.org/0000-0001-6302-6468</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Barna</surname><given-names>Danielle Marie</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-4125-9874</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Engeland</surname><given-names>Kolbjørn</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-1081-8570</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Kolberg</surname><given-names>Sjur Anders</given-names></name>
          
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Nordeide</surname><given-names>Sunniva</given-names></name>
          
        </contrib>
        <aff id="aff1"><label>1</label><institution>Department of Hydrology, Norwegian Water Resources and Energy Directorate, Oslo, Norway</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Sigrid Jørgensen Bakke (sijb@nve.no)</corresp></author-notes><pub-date><day>6</day><month>October</month><year>2026</year></pub-date>
      
      <volume>30</volume>
      <issue>19</issue>
      <fpage>6207</fpage><lpage>6233</lpage>
      <history>
        <date date-type="received"><day>24</day><month>February</month><year>2026</year></date>
           <date date-type="rev-request"><day>6</day><month>March</month><year>2026</year></date>
           <date date-type="rev-recd"><day>26</day><month>August</month><year>2026</year></date>
           <date date-type="accepted"><day>24</day><month>September</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Sigrid Jørgensen Bakke et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026.html">This article is available from https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026.html</self-uri><self-uri xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026.pdf">The full text article is available as a PDF file from https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e116">One of the most important skills of hydrological models is to simulate timing and magnitude of flood events. Long Short-Term Memory (LSTM) networks are currently among the most successful models for streamflow and flood prediction over large regions. In snow-influenced catchments, which typically comprise a minority in large-scale studies, floods are generated by two distinctly different processes, snowmelt and rainfall. The applicability of hydrological models in such regions is therefore dependent on their ability to represent both types of floods. Nevertheless, flood evaluations of LSTM taking different flood-generating processes into account are currently lacking. This study fills this gap by evaluating the ability of LSTM to model flood peak characteristics separately for snowmelt and rainfall generated floods. The trained LSTM model successfully simulated streamflow time series across the 103 evaluated catchments, with average NSE of 0.84 and average KGE of 0.86 over the unseen evaluation period. LSTM exhibited better performance in the majority of the catchments in terms of flood peak timing and magnitude for both rainfall and snowmelt generated floods when compared to the operational hydrological model in the region (HBV) used as a benchmark. LSTM had a 27 pp higher percentage of correctly simulated peak days for rainfall generated floods as compared to snowmelt generated floods, similar to what was found for HBV (29 pp). LSTM's ability to simulate flood peak magnitudes was similar for the different flood types, with percent errors within 20 % for close to half of the events. False alarm rate and probability of detection results were similar for snowmelt and rainfall generated events, both slightly better than those of mixed events. LSTM outperformed HBV in 70 %–91 % of the catchments in terms of flood peak timing and magnitude for snowmelt and rainfall generated floods. The largest improvements in flood peak magnitudes were found for rainfall generated events, in particular for catchments where HBV exhibited high (<inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">40</mml:mn><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="italic">%</mml:mi></mml:mrow></mml:math></inline-formula>) absolute errors. Overall, our findings bring confidence that LSTM can improve hydrological services in regions subject to both snowmelt and rainfall generated floods.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e141">Reliable flood simulations are important for a range of national to local tasks, including area planning, flood and landslide forecasting, hydropower management, and climate change impact assessments. Process-based hydrological models that most commonly perform this task, have recently been challenged by deep learning Long Short-Term Memory (LSTM) networks <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx20" id="paren.1"/>. Variants of LSTM networks have proven their success as rainfall-runoff models across regions, often outperforming benchmark process-based models in terms of streamflow prediction <xref ref-type="bibr" rid="bib1.bibx29" id="paren.2"/>, riverine flood prediction <xref ref-type="bibr" rid="bib1.bibx15" id="paren.3"/> and forecasting <xref ref-type="bibr" rid="bib1.bibx41" id="paren.4"/>. Further, rainfall-runoff modelling by LSTM surpasses other deep learning techniques, such as traditional recurrent neural networks <xref ref-type="bibr" rid="bib1.bibx29" id="paren.5"/>, feedforward neural networks <xref ref-type="bibr" rid="bib1.bibx22" id="paren.6"/>, and Transformer-based models <xref ref-type="bibr" rid="bib1.bibx37" id="paren.7"/>. Unlike most process-based models that perform the best when they are locally calibrated, LSTM, being a purely data-driven model, excels when trained on big datasets of multiple catchments <xref ref-type="bibr" rid="bib1.bibx34 bib1.bibx31" id="paren.8"/>. Thus, LSTM has potential to both improve and simplify hydrological services on local to global scales.</p>
      <p id="d2e169">In mountainous and northern regions, it is particularly important that an applied hydrological model is able to capture effects of seasonal snow accumulation and melt on streamflow. Without physical processes or constraints explicitly implemented in the model code, a purely data-driven model needs to learn such effects solely from the training data. In LSTM, long-term dependencies between input and output time series are handled by memory cells that can represent depletion, increase and outflow of reservoirs and storages in the case of hydrological modelling <xref ref-type="bibr" rid="bib1.bibx30 bib1.bibx29" id="paren.9"/>. Studies have indeed shown how LSTM cell state vectors produce temporal dynamics matching our physical understanding of snow accumulation and melt <xref ref-type="bibr" rid="bib1.bibx29" id="paren.10"/>, and with good (<inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.8</mml:mn></mml:mrow></mml:math></inline-formula>) correlations with snow depth reanalysis products <xref ref-type="bibr" rid="bib1.bibx36" id="paren.11"/>. Correspondingly, LSTMs have demonstrated good performance in snow-influenced regions in terms of simulating streamflow in general <xref ref-type="bibr" rid="bib1.bibx45 bib1.bibx24" id="paren.12"/>, and flood peaks in particular <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx22 bib1.bibx2" id="paren.13"/>.</p>
      <p id="d2e198">As snowmelt and rainfall generated floods are substantially different in their underlying physics and temporal dependencies of past weather, the performance of a hydrological model may differ for the two types of floods. Identifying and understanding such difference are important both for improving models and quantification of uncertainty in model simulations and forecasts. Nevertheless, to our knowledge, existing evaluations of LSTM's ability to simulate floods are performed over all identified flood events, without evaluating events generated by different processes separately. Evaluation results over all flood peaks mainly reflect the model's ability to simulate the flood type possessing the majority of the events, whereas the ability to simulate other flood types of importance is left unknown <xref ref-type="bibr" rid="bib1.bibx39" id="paren.14"/>. This is a particular issue for regions with seasonal snow cover where floods are primarily driven by two fundamentally different processes, snowmelt and rainfall, and an applied hydrological model should be able to simulate both flood generating processes.</p>
      <p id="d2e204">This study meets the abovementioned gap by investigating the ability of LSTM to simulate floods of different driving mechanisms. The main objective is to evaluate LSTM's potential for operational use in snow-influenced regions, focusing on prediction of floods dominantly generated by snowmelt and rainfall separately. The study area is the mainland of Norway, a country spanning 58–71° N that has a large variability in topography, annual precipitation amounts and seasonal snow cover. The applicability of LSTM for a given country or region depends on the performance of LSTM versus the regions' state-of-the-art operational hydrological model in terms of hydrological features that are important in that region. Accordingly, we use a national operational hydrological model as a benchmark. We meet our aim by answering the following research questions: <list list-type="order"><list-item>
      <p id="d2e209">Is LSTM able to simulate streamflow series with similar or exceeding overall performance as compared to the operational model used in the region?</p></list-item><list-item>
      <p id="d2e213">How well does LSTM simulate peak timing and magnitude of all collected floods generated by (i) snowmelt, (ii) rainfall and (iii) a mix of the two?</p></list-item><list-item>
      <p id="d2e217">How well does LSTM simulate per-catchment peak timing and magnitude of (i) snowmelt generated floods and (ii) rainfall generated floods?</p></list-item></list></p>
      <p id="d2e221">The paper is structured as follows: In Sect. <xref ref-type="sec" rid="Ch1.S2"/> the data underlying the analyses are described, followed by methods explaining the applied LSTM and HBV model, flood event detection and classification, and the applied model evaluation metrics. Section <xref ref-type="sec" rid="Ch1.S3"/> presents the results, which are further discussed in Sect. <xref ref-type="sec" rid="Ch1.S4"/> before conclusions are drawn in Sect. <xref ref-type="sec" rid="Ch1.S5"/>.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Data and methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Hydroclimatology and floods in Norway</title>
      <p id="d2e247">Norway exhibits a high heterogeneity in hydroclimatological conditions <xref ref-type="bibr" rid="bib1.bibx12" id="paren.15"/>. The country spans 58 to 71° N along the west side of the Scandinavian peninsula with coastline comprising the southern, western and northern border, and with the Scandinavian mountain range running southwest to northeast. Annual precipitation amounts range from above 6000 mm in the west due to the combination of prevailing westerly winds from the Atlantic and orographic effects, to below 300 mm towards central-eastern and northeastern parts of the country. The geographical pattern of annual runoff reflects to a large degree that of annual precipitation. Annual temperatures varies from <inline-formula><mml:math id="M3" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">9.5</mml:mn></mml:mrow></mml:math></inline-formula> to 9.5 °C, with the warmest areas along the southern and western coastline, and coldest in the inland high altitudes or latitudes. Although the snow season and volumes vary considerably across Norway following the heterogeneity of temperature and precipitation, most of the country has temperatures above 0 °C during summer and below 0 °C and snowfall during winter.</p>
      <p id="d2e263">Floods in Norway are mainly generated by rainfall and snowmelt or a combination of the two <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx55" id="paren.16"/>. Whereas extreme rainfall events typically occur in summer or autumn, snowmelt typically peaks in spring or early summer. Generally, the largest floods are dominantly generated by rainfall in southern and western regions, and by snowmelt in central-eastern and northern regions.</p>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Discharge data and catchment attributes</title>
      <p id="d2e277">Daily discharge data used to train the LSTM network stemmed from 200 streamflow gauging stations of the Norwegian Water Resources and Energy Directorate (NVE) hydrometric observation network in Norway (Fig. <xref ref-type="fig" rid="F1"/>). Of those stations, model evaluation was performed over the 103 stations used by the operational model, which serves as a benchmark (ref. Sect. <xref ref-type="sec" rid="Ch1.S2.SS5"/>). Each day in the observed daily discharge series is divided at midnight Central European Time (CET).</p>

      <fig id="F1"><label>Figure 1</label><caption><p id="d2e286">Locations and catchments of the streamflow gauging stations, and histogram of catchment areas. All 200 catchments (both yellow and dark grey) were used to train the LSTM model, whereas yellow represents the 103 catchments where both LSTM and HBV have simulations (used for model evaluation). Catchment polygons from <xref ref-type="bibr" rid="bib1.bibx17" id="text.17"/>.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f01.png"/>

        </fig>

      <p id="d2e298">The gauging stations were selected from an existing dataset that has been thoroughly quality controlled for flood analyses <xref ref-type="bibr" rid="bib1.bibx13" id="paren.18"><named-content content-type="pre">530 stations;</named-content></xref>. Requirements during that selection included: The stations should not have problems with ice jams or supercritical flow during floods, they should have satisfactory rating curves and measurement conditions during floods, and they should not have homogeneity breaks for streamflow and floods that stemmed from changes in the measurements or quality control. Of the 530 stations, we selected the active stations unaffected by regulations. The original unit of m<sup>3</sup> s<sup>−1</sup> was changed to mm d<sup>−1</sup> by normalising the streamflow values by the catchment areas. Catchment areas represented in the dataset range 0.58 up to 14 000 km<sup>2</sup>. However, most areas are less than 300 km<sup>2</sup>, which was expected due to the requirement of near-natural flow in a country heavily affected by hydropower.</p>
      <p id="d2e359">Table <xref ref-type="table" rid="T1"/> lists the catchment attributes selected for the training of the LSTM model. The 21 catchment attributes were selected to represent the variability of geometric, physiographic and climatological characteristics important for streamflow, hydrological regimes and floods. There is a large degree of overlap between our selection and attributes that have previously been selected for regional flood frequency models in Norway <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx5" id="paren.19"/>. Most of the catchments are influenced by seasonal snow. Of all the 200 catchments, 95 % have average winter (DJF) temperatures below 0 °C (range <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">13</mml:mn></mml:mrow></mml:math></inline-formula> to 3 °C), whereas average summer (JJA) temperatures are all positive, ranging 6 to 16 °C. Several of the catchments are also partly covered by glaciers. Specifically, 21 (12) of the 200 catchments have more than 3 % (10 %) of their catchment area covered by glaciers in 2019.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e380">Catchment attribute names, explanations, units and references to data.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="6cm"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="justify" colwidth="4.5cm"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Name</oasis:entry>
         <oasis:entry colname="col2" align="left">Explanation</oasis:entry>
         <oasis:entry colname="col3">Unit</oasis:entry>
         <oasis:entry colname="col4" align="left">Reference to data</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">area</oasis:entry>
         <oasis:entry colname="col2" align="left">Catchment area</oasis:entry>
         <oasis:entry colname="col3">km<sup>2</sup></oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx17" id="text.20"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">areaperperimeter_km</oasis:entry>
         <oasis:entry colname="col2" align="left">Area divided by circumference</oasis:entry>
         <oasis:entry colname="col3">km</oasis:entry>
         <oasis:entry colname="col4" align="left">Derived from <xref ref-type="bibr" rid="bib1.bibx17" id="text.21"/></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">height_hypso_10</oasis:entry>
         <oasis:entry colname="col2" align="left">10th percentile of hypsographic curve</oasis:entry>
         <oasis:entry colname="col3">masl</oasis:entry>
         <oasis:entry colname="col4" align="left">NVE's database <xref ref-type="bibr" rid="bib1.bibx43" id="paren.22"><named-content content-type="pre">available at</named-content></xref></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">height_hypso_90</oasis:entry>
         <oasis:entry colname="col2" align="left">90th percentile of hypsographic curve</oasis:entry>
         <oasis:entry colname="col3">masl</oasis:entry>
         <oasis:entry colname="col4" align="left">NVE's database <xref ref-type="bibr" rid="bib1.bibx43" id="paren.23"><named-content content-type="pre">available at</named-content></xref></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">length_km_river</oasis:entry>
         <oasis:entry colname="col2" align="left">Length of main river</oasis:entry>
         <oasis:entry colname="col3">km</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx18" id="text.24"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">slope_mean</oasis:entry>
         <oasis:entry colname="col2" align="left">Catchment mean slope</oasis:entry>
         <oasis:entry colname="col3">°</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx25" id="text.25"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">gradient_1085</oasis:entry>
         <oasis:entry colname="col2" align="left">Gradient of main river excluding the 10 % lowest and 15 % highest reaches</oasis:entry>
         <oasis:entry colname="col3">m km<sup>−1</sup></oasis:entry>
         <oasis:entry colname="col4" align="left">NVE's database, based on <xref ref-type="bibr" rid="bib1.bibx18" id="text.26"/></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">drainage_density</oasis:entry>
         <oasis:entry colname="col2" align="left">Drainage density (total river length divided by catchment area)</oasis:entry>
         <oasis:entry colname="col3">km<sup>−1</sup></oasis:entry>
         <oasis:entry colname="col4" align="left">NVE's database, based on <xref ref-type="bibr" rid="bib1.bibx18" id="text.27"/></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">infp</oasis:entry>
         <oasis:entry colname="col2" align="left">Infiltration potential (1–5)</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx19" id="text.28"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">perc_forest</oasis:entry>
         <oasis:entry colname="col2" align="left">Percentage of catchment covered by forest</oasis:entry>
         <oasis:entry colname="col3">%</oasis:entry>
         <oasis:entry colname="col4" align="left">NVE's database <xref ref-type="bibr" rid="bib1.bibx43" id="paren.29"><named-content content-type="pre">available at</named-content></xref></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">perc_eff_lake</oasis:entry>
         <oasis:entry colname="col2" align="left">Effective lake percentage, defined as <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:mn mathvariant="normal">100</mml:mn><mml:mo>⋅</mml:mo><mml:msubsup><mml:mo>∑</mml:mo><mml:mi>l</mml:mi><mml:mi>L</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msub><mml:mi>A</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mspace width="0.25em" linebreak="nobreak"/><mml:msub><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>/</mml:mo><mml:mi>A</mml:mi></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M14" display="inline"><mml:mi>A</mml:mi></mml:math></inline-formula> is catchment area, and <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi>a</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msub><mml:mi>A</mml:mi><mml:mi>l</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are surface area and drainage area of lake <inline-formula><mml:math id="M17" display="inline"><mml:mrow><mml:mi>l</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>L</mml:mi></mml:mrow></mml:math></inline-formula>, respectively</oasis:entry>
         <oasis:entry colname="col3">%</oasis:entry>
         <oasis:entry colname="col4" align="left">NVE's database</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">perc_glacier1985</oasis:entry>
         <oasis:entry colname="col2" align="left">Percentage of catchment covered by glaciers in 1985</oasis:entry>
         <oasis:entry colname="col3">%</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx56" id="text.30"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">perc_glacier2019</oasis:entry>
         <oasis:entry colname="col2" align="left">Percentage of catchment covered by glaciers in 2019</oasis:entry>
         <oasis:entry colname="col3">%</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx3" id="text.31"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">t_djf</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean winter (DJF) temperature 1991–2020</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx53" id="text.32"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">t_mam</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean spring (MAM) temperature 1991–2020</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx53" id="text.33"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">t_jja</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean summer (JJA) temperature 1991–2020</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx53" id="text.34"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">t_son</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean autumn (SON) temperature 1991–2020</oasis:entry>
         <oasis:entry colname="col3">°C</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx53" id="text.35"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">p_djf</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean winter (DJF) precipitation sum 1991–2020</oasis:entry>
         <oasis:entry colname="col3">mm per winter</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx53" id="text.36"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">p_mam</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean spring (MAM) precipitation sum 1991–2020</oasis:entry>
         <oasis:entry colname="col3">mm per spring</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx53" id="text.37"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">p_jja</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean summer (JJA) precipitation sum 1991–2020</oasis:entry>
         <oasis:entry colname="col3">mm per summer</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx53" id="text.38"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">p_son</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean autumn (SON) precipitation sum 1991–2020</oasis:entry>
         <oasis:entry colname="col3">mm per autumn</oasis:entry>
         <oasis:entry colname="col4" align="left">
                      <xref ref-type="bibr" rid="bib1.bibx53" id="text.39"/>
                    </oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Gridded hydrometeorological data</title>
      <p id="d2e919">Daily precipitation sum (mm) and mean daily temperature (°C) from the meteorological dataset SeNorge_2018 <xref ref-type="bibr" rid="bib1.bibx38" id="paren.40"/> was used to produce the dynamical forcing for LSTM (Sect. <xref ref-type="sec" rid="Ch1.S2.SS4"/>). Flood event classification (Sect. <xref ref-type="sec" rid="Ch1.S2.SS6"/>) was based on the same precipitation data, as well as daily snowmelt (mm) estimates simulated by the SeNorge snow model v.1.1.1 <xref ref-type="bibr" rid="bib1.bibx49" id="paren.41"/> that uses SeNorge_2018  gridded daily precipitation sum and mean temperature as input. Both SeNorge_2018 and the snowmelt dataset are available from 1957 to present at a daily time step (07:00–07:00 CET). They are gridded products with a spatial resolution of <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> km<sup>2</sup>, covering the whole of Norway as well as neighbouring regions. The SeNorge_2018 dataset is generated by statistical interpolation of in situ precipitation and temperature observations. The SeNorge snow model used for classifying flood events simulates snowmelt by a degree-day approach augmented with synthetic incoming daily solar irradiance defined as a function of latitude and time of the year.</p>
      <p id="d2e954">From the daily gridded variables precipitation sum and mean temperature, the following daily time series were prepared at the catchment level for LSTM modelling: <list list-type="bullet"><list-item>
      <p id="d2e959">prec (mm d<sup>−1</sup>): Spatially averaged daily precipitation sum</p></list-item><list-item>
      <p id="d2e975">temp_ave (°C): Spatially averaged daily mean temperature</p></list-item><list-item>
      <p id="d2e979">temp_min (°C): The minimum value of all catchment grid cells' daily mean temperature</p></list-item><list-item>
      <p id="d2e983">temp_max (°C): The maximum value of all catchment grid cells' daily mean temperature</p></list-item></list> Further, based on daily gridded precipitation sum and snowmelt, the following time series were prepared at the catchment level to identify flood generating processes: <list list-type="bullet"><list-item>
      <p id="d2e989">snowmelt (mm d<sup>−1</sup>): Spatially averaged daily snowmelt</p></list-item><list-item>
      <p id="d2e1005">rainfall (mm d<sup>−1</sup>): Spatially averaged daily rainfall sum. For each grid cell, the daily rainfall sum was defined as the daily precipitation sum if the daily mean temperature was equal or larger than 0.5 °C, and 0 mm d<sup>−1</sup> otherwise.</p></list-item></list></p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Long Short-Term Memory networks</title>
      <p id="d2e1040">Recurrent neural networks operate on time series data by processing a <inline-formula><mml:math id="M24" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula>-dimensional input sequence <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mi>d</mml:mi></mml:msup><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> and generating a <inline-formula><mml:math id="M26" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-dimensional output sequence <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mi>p</mml:mi></mml:msup><mml:mo>:</mml:mo><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> through recursive computations. At each time step <inline-formula><mml:math id="M28" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, the network updates its internal state and produces an output according to:

            <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M29" display="block"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>f</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mrow><mml:mi mathvariant="normal">t</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mspace linebreak="nobreak" width="1em"/><mml:msub><mml:mi>y</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mi>g</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>h</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>)</mml:mo><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>∈</mml:mo><mml:msup><mml:mi mathvariant="double-struck">R</mml:mi><mml:mi>k</mml:mi></mml:msup></mml:mrow></mml:math></inline-formula> represents the <inline-formula><mml:math id="M31" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-dimensional hidden state that encodes information from previous time steps, while <inline-formula><mml:math id="M32" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M33" display="inline"><mml:mi>g</mml:mi></mml:math></inline-formula> are parametrised functions that define the network's dynamics and output mapping, respectively.</p>
      <p id="d2e1250">LSTM networks implement a sophisticated version of the state update function <inline-formula><mml:math id="M34" display="inline"><mml:mi>f</mml:mi></mml:math></inline-formula> that allows the LSTM to selectively retain or forget information over long sequences, mitigating the vanishing gradient problem that affects traditional recurrent neural networks <xref ref-type="bibr" rid="bib1.bibx20" id="paren.42"/>. In the context of streamflow prediction, this long-term memory capability enables the model to represent snowpack dynamics by retaining information about precipitation and temperature conditions during winter months, then utilizing this stored information to predict the timing and magnitude of snowmelt-driven streamflow in spring and summer <xref ref-type="bibr" rid="bib1.bibx36" id="paren.43"/>. For a detailed decomposition of the LSTM state update function, see <xref ref-type="bibr" rid="bib1.bibx29" id="text.44"/>.</p>
      <p id="d2e1269">To predict streamflow at day <inline-formula><mml:math id="M35" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, the model requires an input sequence <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">t</mml:mi><mml:mo>-</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo><mml:mo>:</mml:mo><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:msub><mml:mi>x</mml:mi><mml:mi mathvariant="italic">τ</mml:mi></mml:msub><mml:mo>:</mml:mo><mml:mi mathvariant="italic">τ</mml:mi><mml:mo>=</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>m</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo>-</mml:mo><mml:mi>m</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula> of length <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:mi>m</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, which is processed through the LSTM to produce a corresponding output sequence <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="normal">t</mml:mi><mml:mo>-</mml:mo><mml:mi>m</mml:mi><mml:mo>)</mml:mo><mml:mo>:</mml:mo><mml:mi mathvariant="normal">t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula>. The output, <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, the streamflow at day <inline-formula><mml:math id="M40" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, is obtained by taking the final element of this output sequence, i.e., <inline-formula><mml:math id="M41" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. To simulate streamflow for the next timestep, the input window is shifted one time step forward at each prediction.</p>
      <p id="d2e1426">We employed LSTM as a lumped rainfall-runoff model using the <monospace>cudalstm</monospace> architecture from the <monospace>neuralhydrology</monospace> package <xref ref-type="bibr" rid="bib1.bibx33" id="paren.45"/>, a Python library containing a collection of LSTM-based models for hydrological modeling built on the PyTorch framework. The <monospace>cudalstm</monospace> model provides an efficient GPU-accelerated implementation of the LSTM architecture described here.</p>
      <p id="d2e1442">The LSTM model was trained using the 15 hydrological years 1 September 2009–31 August 2024, employing the catchment-averaged Nash-Sutcliffe efficiency (NSE) as our loss function, in line with <xref ref-type="bibr" rid="bib1.bibx31" id="text.46"/>. Details of our LSTM training and hyperparameter tuning using a 5-fold cross-validation (CV) are provided in Appendix <xref ref-type="sec" rid="App1.Ch1.S1"/>. The model was trained using four meteorological input time series for each catchment: catchment-averaged daily precipitation sum (mm d<sup>−1</sup>), catchment-averaged mean daily temperature (°C), catchment-minimum mean daily temperature (°C), and catchment-maximum mean daily temperature (°C). In addition to these dynamic inputs, the model incorporated 21 static catchment attributes as listed in Table <xref ref-type="table" rid="T1"/>. Following best practice <xref ref-type="bibr" rid="bib1.bibx34 bib1.bibx31" id="paren.47"/>, a single LSTM model was trained across all catchments, enabling the model to learn general hydrological relationships from the diverse set of catchments.</p>
</sec>
<sec id="Ch1.S2.SS5">
  <label>2.5</label><title>Benchmark model</title>
      <p id="d2e1476">The Hydrologiska Byråns Vattenbalansavdelning (HBV) model <xref ref-type="bibr" rid="bib1.bibx8" id="paren.48"/> is a widely used precipitation-runoff model, in particularly in the Nordic countries <xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx50" id="paren.49"/>. In Norway, a version of the HBV model with a daily temporal resolution is used operationally by NVE, for national flood warning services, quality control of streamflow observations, and energy prognosis among others <xref ref-type="bibr" rid="bib1.bibx35 bib1.bibx46" id="paren.50"/>. For this reason, we chose the operational HBV model for benchmarking our results. Catchment-by-catchment calibration of 12 model parameters were performed using Nash-Sutcliffe efficiency (NSE) as a loss function (for calibration details, see Appendix <xref ref-type="sec" rid="App1.Ch1.S2"/>). The operational HBV model are currently used for 145 reference catchments with long data series of good quality and limited influence of regulations, and representative for different geographical regions in Norway <xref ref-type="bibr" rid="bib1.bibx46" id="paren.51"/>. We performed all model evaluations for the 103 reference catchments that overlapped with the catchments used for training LSTM (Fig. <xref ref-type="fig" rid="F1"/>).</p>
</sec>
<sec id="Ch1.S2.SS6">
  <label>2.6</label><title>Detection of flood events and flood generating processes</title>
      <p id="d2e1504">Flood events were selected from the observational time series based on a peak over threshold (POT) approach, following the methods described in <xref ref-type="bibr" rid="bib1.bibx55" id="text.52"/>. Daily streamflow values exceeding the 98th percentile calculated from the 30-year period 1 September 1994–31 August 2024 were used to identify the flood peaks. To ensure independence between flood events, only the maximum streamflow value was selected within a catchment-specific time window, using the python library <monospace>pyextremes</monospace> <xref ref-type="bibr" rid="bib1.bibx9" id="paren.53"/>.  This time window was set to twice the normal flood duration (NFD), which is defined as the sum of concentration time plus recession time of the largest flood events. The full description of the estimation of NFD is provided in <xref ref-type="bibr" rid="bib1.bibx55" id="text.54"/>. In general, a concentration time of 2–3 d was found for all catchments subject to an HBV modelling experiment inputting twice the annual rainfall over a 2 d period to a fully drained catchment. The median of two days was then chosen for all catchments. Recession times were estimated using the methods described in <xref ref-type="bibr" rid="bib1.bibx51" id="text.55"/>, which fit a recession curve to the largest flood events from the observational data series. Our 200 gauging stations have NDF spanning 2 to 45 d.</p>
      <p id="d2e1522">Flood generating process (FGP) was defined for each event based on the relative contribution of rainfall and snowmelt to the flood, based on the approach in <xref ref-type="bibr" rid="bib1.bibx54" id="text.56"/>. We selected this approach from the range of available flood classification schemes <xref ref-type="bibr" rid="bib1.bibx52" id="paren.57"/> because it provides a simple yet effective distinction between rainfall and snowmelt generated floods and has been successfully applied in studies conducted within our region <xref ref-type="bibr" rid="bib1.bibx58 bib1.bibx55 bib1.bibx54" id="paren.58"/>. To account for antecedent conditions, the contribution to a specific flood was computed as the sum of rainfall and snowmelt in a catchment-specific time window prior to, and including, the flood peak day. This time window was set as the catchment-specific recession times as defined above. Figure <xref ref-type="fig" rid="F2"/> shows the average fractional contribution of rainfall to flood events, reflecting the dominant flood generating process (FGP) in each catchment. Using the thresholds proposed by <xref ref-type="bibr" rid="bib1.bibx54" id="text.59"/>, all flood events were classified into snowmelt generated events, rainfall generated events and mixed events depending on whether more than two thirds of the contribution stemmed from snowmelt, rainfall or neither, respectively. The distribution of all flood events' fractional rainfall contribution, along with the flood type classification thresholds, is shown in Fig. <xref ref-type="fig" rid="FC3"/>. The figure further shows two alternative pairs of classification thresholds used to assess the results' sensitivity to the chosen thresholds.</p>

      <fig id="F2"><label>Figure 2</label><caption><p id="d2e1544">Average fractional rainfall (i.e. proportion of total) contribution to all observed flood events in the period 1 September 1994–31 August 2024, and its relation to the catchment dominant flood generating process (FGP). Dark pink (green) dots represent catchments with snowmelt (rainfall) as their dominant FGP, whereas near-white dots represent catchments with similar average contributions from rainfall and snowmelt. All 200 catchments used for training the LSTM are shown, with thicker circles representing the 103 catchments used for model evaluation.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f02.png"/>

        </fig>

</sec>
<sec id="Ch1.S2.SS7">
  <label>2.7</label><title>Model evaluation</title>
      <p id="d2e1561">Both LSTM and our benchmark HBV model were evaluated using a period unseen during model training/calibration, i.e. 15 hydrological years comprising 1 September 1994–31 August 2009. Early 1990s is typically considered the start of thoroughly quality-controlled rating curves for a large proportion of the Norwegian streamflow gauging stations. We did not want the models to suffer from changes in the water balance from the calibration period to the evaluation period due to changes in the precipitation measuring network. To assess whether this was the case, we computed the runoff coefficient, defined as the ratio of streamflow to precipitation, for the two periods. For most catchments, the runoff coefficients were similar between the two periods (Fig. <xref ref-type="fig" rid="FC1"/>). However, for four (13) catchments, a 30 % (20 %) change in the runoff coefficient from training to evaluation period was detected, which might have had an influence on the model performance results. Negative discharge values in the LSTM simulations were truncated to zero prior to the evaluation, and simulated series for each catchment were masked by data gaps existing in the observed series.</p>
      <p id="d2e1566">We evaluated model performance in terms of (i) per-catchment overall performance, (ii) descriptive statistics over all collected peaks of each flood type, and (iii) per-catchment flood peak performance of each flood type. In terms of (i), we assessed the fit between observed and simulated streamflow across the evaluation period for each catchment using Nash-Sutcliffe efficiency (NSE), the updated Kling-Gupta efficiency <xref ref-type="bibr" rid="bib1.bibx28" id="paren.60"><named-content content-type="pre">KGE;</named-content></xref>, and mean absolute error (MAE), all defined in Table <xref ref-type="table" rid="T2"/>. Additionally, we assessed the three components comprising KGE: Pearson correlation coefficient (<inline-formula><mml:math id="M43" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula>), bias ratio (<inline-formula><mml:math id="M44" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>) and variability ratio (<inline-formula><mml:math id="M45" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>). The two latter are defined as:

            <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M46" display="block"><mml:mrow><mml:mi mathvariant="italic">β</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mover accent="true"><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

          and

            <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M47" display="block"><mml:mrow><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover></mml:msub><mml:mo>/</mml:mo><mml:mover accent="true"><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>q</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

          where <inline-formula><mml:math id="M48" display="inline"><mml:mover accent="true"><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> is mean simulated streamflow, <inline-formula><mml:math id="M49" display="inline"><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:math></inline-formula> is mean observed streamflow, <inline-formula><mml:math id="M50" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover></mml:msub></mml:mrow></mml:math></inline-formula> is standard deviation of simulated streamflow, and <inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>q</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is standard deviation of observed streamflow.</p>
      <p id="d2e1715">In step (ii), we used descriptive statistics to summarise the flood peak performance in terms of timing and magnitudes over all collected flood peaks (i.e. independent of catchment) of each flood type: snowmelt generated floods, mixed floods and rainfall generated floods. Timing was evaluated by computing the percentage of peaks of a given flood type simulated at the day of observed peak (i.e. correct timing), one day too early or too late, or more than one day too early or too late. Flood peak magnitudes were evaluated by percent errors, considering both the simulated discharge magnitude at the day of observed flood peak (<inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) and the simulated maximum discharge within a 2 d time window of the observed flood peak (<inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). The latter allows comparison of peak magnitudes even if the model does not match the exact day of the observed peak. The definitions of these metrics are as follows. For each catchment <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:mi>s</mml:mi><mml:mo>∈</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:math></inline-formula>, let <inline-formula><mml:math id="M55" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> be the number of observed floods of a specific type at catchment <inline-formula><mml:math id="M56" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula>. Let <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denote the timestep of the <inline-formula><mml:math id="M58" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th observed flood peak. Define <inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>:=</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>:=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> as the observed and simulated streamflow at the <inline-formula><mml:math id="M61" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>th peak, respectively. Then, <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> for peak <inline-formula><mml:math id="M63" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> at catchment <inline-formula><mml:math id="M64" display="inline"><mml:mi>s</mml:mi></mml:math></inline-formula> is

            <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M65" display="block"><mml:mrow><mml:msubsup><mml:mtext>PE</mml:mtext><mml:mtext>peak,pinned</mml:mtext><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>⋅</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

          for <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M67" display="inline"><mml:mrow><mml:mi>s</mml:mi><mml:mo>∈</mml:mo><mml:mi>S</mml:mi></mml:mrow></mml:math></inline-formula>. Further, defining the window <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:msub><mml:mi>I</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mo mathvariant="italic">{</mml:mo><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:mi>Z</mml:mi><mml:mo>:</mml:mo><mml:mo>|</mml:mo><mml:mi>i</mml:mi><mml:mo>-</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:mo>≤</mml:mo><mml:mn mathvariant="normal">2</mml:mn><mml:mo mathvariant="italic">}</mml:mo></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M69" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> is

            <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M70" display="block"><mml:mrow><mml:msubsup><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mo>max⁡</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>∈</mml:mo><mml:msub><mml:mi>I</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mrow><mml:mi>s</mml:mi><mml:mo>,</mml:mo><mml:mi>j</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>⋅</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

          Further, we computed the false alarm rate (FAR) and probability of detection (POD) of each flood type <xref ref-type="bibr" rid="bib1.bibx57" id="paren.61"><named-content content-type="pre">e.g.</named-content></xref>. FAR measures the percentage of false positives, i.e. incorrectly identified flood events by a given model, and has an optimum value at 0 %. To compute FAR, we applied the flood event selection and classification procedure described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS6"/> on the simulated streamflow series. For each detected simulated flood event, the corresponding “pinned” (i.e. exact day) and “floating” (i.e. maximum within a 2 d time window) observed streamflow value was extracted and defined as a non-flood event if the value did not exceed the 98th percentile. <inline-formula><mml:math id="M71" display="inline"><mml:mrow><mml:msub><mml:mtext>FAR</mml:mtext><mml:mtext>pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M72" display="inline"><mml:mrow><mml:msub><mml:mtext>FAR</mml:mtext><mml:mtext>floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> were then defined as the percent of simulated flood events that had corresponding pinned, respectively floating, observed non-flood events. POD measures to what degree a model is able to capture the observed flood events. It has an optimum at 100 %. <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:msub><mml:mtext>POD</mml:mtext><mml:mtext>pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mtext>POD</mml:mtext><mml:mtext>floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> were computed as the percent of observed flood events that had a corresponding pinned, respectively floating, simulated streamflow exceeding the 98th percentile. We compared FAR and POD results using the 98th percentile streamflow magnitudes extracted from each individual (simulated and observed) time series with results using only the observed series to define the 98th percentile streamflow magnitudes.</p>

<table-wrap id="T2" specific-use="star"><label>Table 2</label><caption><p id="d2e2172">Definitions of per-catchment evaluation metrics used in this study: Nash-Sutcliffe efficiency (NSE), Kling-Gupta efficiency (KGE), mean absolute error (MAE), mean absolute percentage error (<inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) and timing error. Here <inline-formula><mml:math id="M77" display="inline"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are observed and simulated streamflow for a specific catchment at timestep <inline-formula><mml:math id="M79" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M80" display="inline"><mml:mrow><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>T</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:msub><mml:mi>q</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the mean observed streamflow, <inline-formula><mml:math id="M81" display="inline"><mml:mrow><mml:mi mathvariant="double-struck">I</mml:mi><mml:mo>(</mml:mo><mml:mo>⋅</mml:mo><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> is the indicator function and <inline-formula><mml:math id="M82" display="inline"><mml:mrow><mml:msub><mml:mi>d</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>t</mml:mi><mml:mo mathvariant="normal" stretchy="false">^</mml:mo></mml:mover><mml:mi>j</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>t</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the difference in days between the simulated and observed peak for event <inline-formula><mml:math id="M83" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula> <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>n</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M85" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is the number of observed floods of a specific type at that catchment. We used adjusted KGE such that <inline-formula><mml:math id="M86" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula> is the Pearson correlation coefficient, <inline-formula><mml:math id="M87" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula> is the bias ratio (Eq. <xref ref-type="disp-formula" rid="Ch1.E2"/>), and <inline-formula><mml:math id="M88" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula> is the variability ratio (Eq. <xref ref-type="disp-formula" rid="Ch1.E3"/>), all computed over the period <inline-formula><mml:math id="M89" display="inline"><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">…</mml:mi><mml:mo>,</mml:mo><mml:mi>T</mml:mi></mml:mrow></mml:math></inline-formula>. The quantities <inline-formula><mml:math id="M90" display="inline"><mml:mrow><mml:msubsup><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext><mml:mi>j</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M91" display="inline"><mml:mrow><mml:msubsup><mml:mtext>PE</mml:mtext><mml:mtext>peak,pinned</mml:mtext><mml:mi>j</mml:mi></mml:msubsup></mml:mrow></mml:math></inline-formula> are defined in Eqs. (<xref ref-type="disp-formula" rid="Ch1.E4"/>) and (<xref ref-type="disp-formula" rid="Ch1.E5"/>), respectively.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="justify" colwidth="4cm"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Evaluation metric</oasis:entry>
         <oasis:entry colname="col2" align="left"/>
         <oasis:entry colname="col3"/>
         <oasis:entry colname="col4">Unit</oasis:entry>
         <oasis:entry colname="col5">Range</oasis:entry>
         <oasis:entry colname="col6">Optimum</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col6">Per-catchment overall performance (over the full time series) </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">NSE</oasis:entry>
         <oasis:entry colname="col2" align="left">Nash-Sutcliffe efficiency</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M92" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msub><mml:mi>q</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">1</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">KGE</oasis:entry>
         <oasis:entry colname="col2" align="left">Kling-Gupta efficiency</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mrow><mml:mo>(</mml:mo><mml:mi mathvariant="italic">ρ</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mi mathvariant="italic">β</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mo>-</mml:mo><mml:mi mathvariant="normal">∞</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">1</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">MAE</oasis:entry>
         <oasis:entry colname="col2" align="left">Mean absolute error</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>T</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>T</mml:mi></mml:msubsup><mml:mfenced open="|" close="|"><mml:mrow><mml:msub><mml:mi>q</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>q</mml:mi><mml:mo stretchy="false" mathvariant="normal">^</mml:mo></mml:mover><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">mm d<sup>−1</sup></oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">∞</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry namest="col1" nameend="col6">Per-catchment flood peak performance for a specific flood type </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Timing error</oasis:entry>
         <oasis:entry colname="col2" align="left">Percent flood peaks simulated at a different day than the corresponding observed flood peaks</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mi mathvariant="double-struck">I</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>d</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>≠</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">%</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">100</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">MAPE<sub>peak,pinned</sub></oasis:entry>
         <oasis:entry colname="col2" align="left">Mean absolute percentage error of simulated discharge at the day of observed flood peak (“pinned”)</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mfenced open="|" close="|"><mml:mrow><mml:msubsup><mml:mtext>PE</mml:mtext><mml:mtext>peak,pinned</mml:mtext><mml:mi>j</mml:mi></mml:msubsup></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">%</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">∞</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">MAPE<sub>peak,floating</sub></oasis:entry>
         <oasis:entry colname="col2" align="left">Mean absolute percentage error of simulated maximum discharge within a 2 d time window of the observed flood peak (“floating”)</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>j</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mfenced open="|" close="|"><mml:mrow><mml:msubsup><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext><mml:mi>j</mml:mi></mml:msubsup></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">%</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">∞</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col6">0</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3000">In the final evaluation step (iii), we evaluated per-catchment timing error and mean absolute percentage error of flood peaks separately for each flood type, as defined in Table <xref ref-type="table" rid="T2"/>. Timing error represents the percent of flood peaks that are simulated at a different day than the observed flood peaks. Corresponding to the definitions of <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> above, we computed the mean absolute percentage error of flood peaks considering both pinned (<inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) and floating (<inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) simulated peaks. The per-catchment flood peak performance metrics were only computed for catchments with minimum ten flood events of the given flood type in the evaluation period. A minimum of ten flood events during the evaluation period were found for 45 catchments for snowmelt generated floods, 60 catchments for rainfall generated floods (including 13 of the same catchments that had minimum ten snowmelt generated floods), and five catchments for mixed floods. Due to the low number of catchments with the required number of mixed flood events, the results for mixed floods are excluded from the per-catchment flood peak performance results.</p>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Results</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Overall model performance</title>
      <p id="d2e3065">The LSTM model performance in terms of NSE and KGE are shown for the evaluated catchments in Fig. <xref ref-type="fig" rid="F3"/>. Importantly, all results were computed for the unseen evaluation period (ref. Sect. <xref ref-type="sec" rid="Ch1.S2.SS7"/>). Average scores were 0.84 for NSE and 0.86 for KGE, and NSE (KGE) scores exceeded 0.7 in 100 (102) catchments. The KGE components (<inline-formula><mml:math id="M111" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M112" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M113" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>) and MAE results are presented in Fig <xref ref-type="fig" rid="FC2"/>. All <inline-formula><mml:math id="M114" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula> (i.e. correlation coefficient) scores exceeded 0.83, whereof 83 % of them exceeded 0.9. A total of 82 % of the catchments had a <inline-formula><mml:math id="M115" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula> (i.e. bias ratio) in the range 0.9–1.1 with a similar percentage of catchments below (46 %) and above (54 %) the optimum of 1. For <inline-formula><mml:math id="M116" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula> (variability ratio), 70 % of the catchments had values in the range 0.9 to 1.1, and in 82 % of the catchments the coefficient of variation was underestimated (i.e. <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:mi mathvariant="italic">γ</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>). MAE scores ranged 0.2 to 4.2 mm d<sup>−1</sup> with a average of 1.1 mm d<sup>−1</sup>.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e3156"><bold>(a)</bold> Nash-Sutcliffe efficiency (NSE) and <bold>(b)</bold> Kling-Gupta efficiency (KGE) of LSTM for the 103 evaluated catchments.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f03.png"/>

        </fig>

      <p id="d2e3170">To benchmark our results, the LSTM NSE and KGE values were compared with those of the HBV model (Fig. <xref ref-type="fig" rid="F4"/>). A total of 99 catchments (96 %) had a higher NSE score for LSTM as compared to HBV, with the score difference exceeding 0.1 for 27 of the catchments. In terms of KGE, a higher LSTM score was found for 79 of the 103 catchments (77 %), with a KGE difference exceeding 0.1 for 12 catchments. One of the HBV KGE scores exceeded the corresponding LSTM scores by more than 0.1. The overweight of higher LSTM scores is reflected in the overweight of catchments above the diagonal line in Fig. <xref ref-type="fig" rid="F4"/>c, d. The largest improvement in scores when comparing LSTM to HBV were found for catchments with HBV scores below 0.6. Catchments with floods predominantly generated by snowmelt (i.e. pink coloured dots) had generally higher NSE scores than catchments predominantly generated by rainfall (i.e. green coloured dots) for HBV. A similar pattern was not found for LSTM's NSE scores, nor for any of the models' KGE scores. In terms of the KGE components, LSTM had a better score for <inline-formula><mml:math id="M120" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula> in 93 %, <inline-formula><mml:math id="M121" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula> in 74 %, and <inline-formula><mml:math id="M122" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula> in 44 % of the catchments. In all but three catchments, mean absolute errors were lower (i.e. better) for LSTM as compared to HBV.</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e3201">Difference between LSTM and HBV in <bold>(a)</bold> NSE and <bold>(b)</bold> KGE. Blue color implies a higher score for LSTM for that catchment, whereas red color implies a higher score for HBV. Scatterplots show <bold>(c)</bold> NSE and <bold>(d)</bold> KGE for HBV (<inline-formula><mml:math id="M123" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) versus LSTM (<inline-formula><mml:math id="M124" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis). Background is shaded by the corresponding 2D empirical density function. Each dot represents a catchment, and dots above the diagonal line represent higher scores for LSTM as compared to HBV. Triangles with adjacent numbers represent catchment average score for each model. Each catchment is coloured by the corresponding catchment's dominant flood generating process (FGP, ref. Fig. <xref ref-type="fig" rid="F2"/>). Inserted in the lower left corner of <bold>(c)</bold> and <bold>(d)</bold> are the cumulative density functions (CDFs) of the two models' NSEs and KGEs, respectively.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f04.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Model performance over all peaks of each flood type</title>
      <p id="d2e3253">Figure <xref ref-type="fig" rid="F5"/> shows the overall ability of LSTM and HBV to simulate the correct timing of the flood peaks, shown separately for snowmelt generated events, mixed events and rainfall generated events. Both models had a notably higher percentage of correctly simulated peak day for rainfall generated events as compared to mixed events, and mixed events as compared to snowmelt generated events. LSTM simulated the correct peak day for 50 % of the snowmelt generated events, 67 % of the mixed events and 77 % of the rainfall generated event. The corresponding correct peak day percentages for HBV were 14–16 pp smaller. For all flood types, a higher percentage of peak events were simulated too early by HBV than by LSTM. For example, 40 % of the snowmelt generated events are simulated too early by HBV, as compared to 21 % for LSTM. For rainfall generated events, the corresponding numbers are 27 % for HBV and 9 % for LSTM. The corresponding results using two alternative pairs of flood type classification thresholds were very similar (Figs. <xref ref-type="fig" rid="FC4"/>–<xref ref-type="fig" rid="FC5"/>).</p>

      <fig id="F5"><label>Figure 5</label><caption><p id="d2e3264">Flood peak timing results for <bold>(a)</bold> LSTM and <bold>(b)</bold> HBV of all collected snowmelt generated floods (“MELT”), mixed floods (“MIXED”) and rainfall generated floods (“RAIN”). Shown are the percentages of flood events where the peak day was simulated at the correct day (green), one day too early (light grey), even earlier (dark grey), one day too late (light purple) or even later (dark purple) as compared to observed flood peak day. <inline-formula><mml:math id="M125" display="inline"><mml:mi>N</mml:mi></mml:math></inline-formula> values in parenthesis represent the number of flood events.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f05.png"/>

        </fig>

      <p id="d2e3286">The distributions of percent error of simulated flood peak magnitudes are relatively similar for snowmelt, mixed and rainfall generated events (Fig. <xref ref-type="fig" rid="F6"/>). Hence, the two alternative pairs of flood type classification thresholds tested produced very similar results (Figs. <xref ref-type="fig" rid="FC6"/>–<xref ref-type="fig" rid="FC7"/>). Slightly better results were found for <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> as compared to <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, as expected since the former allows for extracting the simulated flood peak despite the model missing the peak day by 1 or 2 d (ref. Fig. <xref ref-type="fig" rid="F5"/>). For LSTM, close to half of the snowmelt (46 %) and rainfall generated events (49 %) are within <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % of <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> values, compared to 36 %, respectively 38 %, of the events for HBV. Regardless of model and flood type, most (77 % to 91 %) of the observed peak magnitudes were underestimated. More events had relatively large underestimations (<inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">40</mml:mn></mml:mrow></mml:math></inline-formula> % to <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> %) for HBV as compared to LSTM.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e3364">Percent error (PE) of flood peak magnitudes simulated by HBV (orange) and LSTM (blue) for the three different types of flood events: snowmelt generated floods <bold>(a, d)</bold>, mixed floods <bold>(b, e)</bold> and rainfall generated floods <bold>(c, f)</bold>. Upper panel <bold>(a–c)</bold> shows the pinned simulations (<inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> comparing discharge magnitudes at the day of observed peak), whereas lower panel <bold>(d–f)</bold> shows the floating simulations (<inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:msub><mml:mtext>PE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> comparing observed peak discharge with maximum simulated discharge within 2 d of observed peak). In each subplot, the percentage of events that were underestimated (numbers at horizontal line extending <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">100</mml:mn></mml:mrow></mml:math></inline-formula> % to 0 %), and the percentage of events with a relative error within <inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % (numbers at horizontal line extending <inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % to 20 %) are shown for LSTM (blue) and HBV (orange).</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f06.png"/>

        </fig>

      <p id="d2e3441">FAR and POD results for the different flood types and models are shown in Fig. <xref ref-type="fig" rid="F7"/>. For FAR, a lower value is desirable, whereas a higher value is desirable for POD. As expected, all results allowing for a two day shift in the flood peak timing (“floating”) are better than corresponding results requiring exact day match (“pinned”). Further, using time series specific flood thresholds generally resulted in higher percentages than corresponding results using only observed flood thresholds. LSTM results ranged 21 %–39 % for FAR and 52 %–70 % for POD. LSTM's FAR and POD results for rainfall and snowmelt generated flood events were generally similar, and 6–11 pp better than those of mixed flood events. HBV had slightly higher FAR values (except <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:msub><mml:mtext>FAR</mml:mtext><mml:mtext>floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> for mixed and rainfall generated events) and lower POD values than LSTM.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e3459">False alarm rate (FAR) and probability of detection (POD) results for LSTM and HBV of all collected snowmelt generated floods (“MELT”), mixed floods (“MIXED”) and rainfall generated floods (“RAIN”). FAR (optimum at 0 %) represents the percent of simulated flood events not exceeding the 98th percentile flood threshold in the observed series <bold>(a)</bold> at the same day (<inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msub><mml:mtext>FAR</mml:mtext><mml:mtext>pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) or <bold>(b)</bold> within 2 d of simulated peak (<inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:msub><mml:mtext>FAR</mml:mtext><mml:mtext>floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). POD (optimum at 100 %) represents the percent of observed flood events that also exceed the 98th percentile flood threshold in the simulated series <bold>(c)</bold> at the same day (<inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:msub><mml:mtext>POD</mml:mtext><mml:mtext>pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) or <bold>(d)</bold> within 2 d of observed peak (<inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msub><mml:mtext>POD</mml:mtext><mml:mtext>floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). Tallest bars represent results when using the 98th percentile discharge magnitudes defined from each corresponding (simulated or observed) time series, whereas the horizontal white lines represent results when only the observed 98th percentile discharge magnitudes were used.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f07.png"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Per-catchment flood peak performance</title>
      <p id="d2e3533">The resulting per-catchment flood peak performance scores for LSTM and HBV are shown in Fig. <xref ref-type="fig" rid="F8"/>, and maps showing the differences in scores between the two models are presented in Fig. <xref ref-type="fig" rid="F9"/>. Across all peak metrics and flood types, LSTM had better scores than HBV for the majority of the catchments. LSTM had the lowest percentage of timing error in 84 % of the catchments for snowmelt generated flood events, and 78 % of the catchments in terms of rainfall generated flood events. Median improvement by LSTM across catchments was 14 pp for snowmelt and 11 pp for rainfall generated floods, but improvements exceeding 30 pp were found for several catchments. Both models had a markedly lower timing error when considering rainfall generated events (catchment averages of 24 % for LSTM and 38 % for HBV) as compared to snowmelt generated events (catchment averages of 49 % for LSTM and 66 % for HBV).</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e3542">Flood peak performance scores per catchment for HBV (<inline-formula><mml:math id="M142" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) versus LSTM (<inline-formula><mml:math id="M143" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis), for snowmelt generated floods (left) and rainfall generated floods (right). Upper panel <bold>(a–b)</bold> shows the peak timing error (percentages of incorrectly simulated days of peak discharges). Middle panel <bold>(c–d)</bold> and lower panel <bold>(e–f)</bold> show the mean absolute percentage error for pinned simulations (<inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> comparing discharge magnitudes at the day of observed peak), and for floating simulations (<inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> comparing observed peak discharge with maximum simulated discharge within two days of observed peak), respectively. Background is shaded by the corresponding 2D empirical density function. Triangles with adjacent numbers represent catchment average score for each model. Dots below the diagonal line represent catchments with better (i.e. smaller error) metric values for LSTM as compared to HBV. The percentages of catchments above and below the diagonal lines are provided in the upper right corner of each scatterplot. Each catchment is coloured by the corresponding catchment's dominant flood generating process (FGP, ref. Fig. <xref ref-type="fig" rid="F2"/>).</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f08.png"/>

        </fig>

      <fig id="F9" specific-use="star"><label>Figure 9</label><caption><p id="d2e3601">Model difference (LSTM minus HBV results) in per-catchment flood peak performance scores for snowmelt generated floods (left) and rainfall generated floods (right). Upper panel <bold>(a–b)</bold> shows the model difference in timing error. Middle panel <bold>(c–d)</bold> and lower panel <bold>(e–f)</bold> show the model difference in mean absolute percentage error for pinned simulations (<inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> comparing discharge magnitudes at the day of observed peak), and for floating simulations (<inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> comparing observed peak discharge with maximum simulated discharge within two days of observed peak), respectively. Blue colour represents a smaller error for LSTM as compared to HBV. Alongside each map, the corresponding histogram, range, median and average difference are provided.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f09.png"/>

        </fig>

      <p id="d2e3642">Most <inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> scores were in the range 10 % to 40 %, with catchment average of 25 % to 26 % for LSTM and 32 % for HBV. <inline-formula><mml:math id="M149" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> for snowmelt generated floods had the largest median model difference of 6.5 pp, in the favour of LSTM. For both error metrics considering peak magnitude (<inline-formula><mml:math id="M150" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>), LSTM had smaller errors than HBV for a larger proportion of the catchments when considering snowmelt generated floods as compared to rainfall generated floods. On the other hand, more catchments had a relatively larger difference between the models when considering rainfall generated floods, in particular for catchments where HBV had a relatively high (<inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">40</mml:mn></mml:mrow></mml:math></inline-formula> %) error score.</p>
      <p id="d2e3699">Figure <xref ref-type="fig" rid="F10"/> combines the model comparisons of overall performance (KGE) and flood peak performance (<inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). Catchments in quadrant IV (lower right square in each plot) have a better metric score for LSTM as compared to HBV in terms of both KGE and <inline-formula><mml:math id="M154" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>. The majority of catchments (76 % and 75 %) are found in quadrant IV. A few catchments show notably better results for LSTM, with a <inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula> larger KGE and a smaller <inline-formula><mml:math id="M156" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> of more than 5 pp. The minority (2 % and 8 %) of the catchments are found in quadrant II that represents better HBV scores for both metrics. Figures <xref ref-type="fig" rid="FC8"/> and <xref ref-type="fig" rid="FC9"/> show model difference in NSE, respectively KGE, versus model difference in all three flood peak metrics. In the corresponding plots using NSE instead of KGE, nearly all catchments are located in quadrants I and IV, as LSTM had the highest NSE scores for 96 % of the catchments.</p>

      <fig id="F10" specific-use="star"><label>Figure 10</label><caption><p id="d2e3754">Model difference (i.e. LSTM minus HBV) in the overall performance metric KGE (<inline-formula><mml:math id="M157" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) versus the peak metric <inline-formula><mml:math id="M158" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> (<inline-formula><mml:math id="M159" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis) for <bold>(a)</bold> snowmelt generated floods and <bold>(b)</bold> rainfall generated floods. Each dot represents a catchment and is coloured by the catchment dominant flood generating process (FGP; ref. Fig. <xref ref-type="fig" rid="F2"/>). Catchments within quadrant IV (QIV; lower right square) have a better score for LSTM as compared to HBV both in terms of overall performance score and flood peak performance score. Percentages of catchments within each quadrant are given in the corners of the plots.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f10.png"/>

        </fig>

      <p id="d2e3797">A per-catchment summary of the best-performing model in terms of all evaluated metrics is presented in Fig. <xref ref-type="fig" rid="F11"/>. LSTM had the best scores for the majority of the catchments for all evaluated metrics except (the KGE component) <inline-formula><mml:math id="M160" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>. For MAE, NSE and <inline-formula><mml:math id="M161" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula>, 93 % to 97 % of the catchments had a better score for LSTM. To indicate the catchment specific best model across metrics, the percentage of metrics where LSTM had a better score is given at the top of each catchment column. For eight of the 103 catchments, HBV had better scores for minimum two thirds of the metrics. LSTM on the other hand, had better scores for minimum two thirds of the metrics for 90 of the catchments.</p>

      <fig id="F11" specific-use="star"><label>Figure 11</label><caption><p id="d2e3818">Summary of best performing model in terms of all evaluated metrics and catchments. LSTM has the best score for blue cells, HBV has the best score for light grey cells, and both models have equal scores for dark grey cells. Percentages of blue cells for each catchment (i.e. column) and metric (i.e. row) are provided at the top and right, respectively.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f11.png"/>

        </fig>

</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Discussion</title>
      <p id="d2e3836">The LSTM model demonstrated good performance across the metrics assessed in the study. Whereas flood peak timing results were notably better for rainfall as compared to snowmelt generated floods, flood peak magnitudes were simulated with comparable performance for different types of floods. As compared to the benchmark model (HBV), performances were improved for a majority of the catchments in terms of all but one evaluated metrics. The largest improvements by LSTM were often found for metrics and catchments where HBV scores were relatively poor. Thus, the results imply that LSTM can improve hydrological services and flood specific assessments in regions affected by both snowmelt and rainfall generated floods.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>LSTM's ability to simulate different aspects of the streamflow series</title>
      <p id="d2e3846">LSTM NSE scores were generally high, exceeding 0.7 for 100 catchments, and with an average NSE of 0.84 as compared to 0.76 for HBV. For one catchment in particular (catchment ID 2.439), the models were unable represent the time series well, with <inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:mi mathvariant="normal">NSE</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.6</mml:mn></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:mi mathvariant="normal">KGE</mml:mi><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula> for both LSTM and HBV. This catchment had calibration NSE of 0.90 (LSTM) and 0.79 (HBV), and a 30 % drop in runoff coefficient from training to evaluation period, indicating problems with the observed meteorological or streamflow data during the evaluation period rather than an inability of the models to represent the catchment's streamflow behaviour. Many of the snowmelt dominated catchments were among the catchments with the highest NSE scores for HBV, in line with previous studies demonstrating that models typically achieve higher NSE and KGE values in catchments with strong seasonality <xref ref-type="bibr" rid="bib1.bibx47 bib1.bibx2" id="paren.62"/>. However, for LSTM in terms of NSE, and for both models in terms of KGE, there were no distinct relation between catchment dominant FGP and model performance score.</p>
      <p id="d2e3876">A higher percentage of catchments were improved by LSTM in terms of NSE (96 %) as compared to KGE (77 %). Of the three components constituting KGE, the one reflecting the variability (<inline-formula><mml:math id="M164" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>) of the model was the only evaluated metric with a better HBV score for the majority of the catchments (Fig. <xref ref-type="fig" rid="F11"/>). LSTM generally underestimated <inline-formula><mml:math id="M165" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula> for more catchments and to a larger degree than HBV (Fig. <xref ref-type="fig" rid="FC2"/>). The underestimated variability may partly relate to the applied loss function (NSE) which normalises the prediction errors by the variability of the time series. <xref ref-type="bibr" rid="bib1.bibx21" id="text.63"/> has previously shown how the variability has to be underestimated to maximise NSE. In terms of the bias ratio (<inline-formula><mml:math id="M166" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>), the majority of catchments exhibited better scores for LSTM than HBV. Due to precipitation undercatch, the HBV model includes catchment specific correction factors for rain and snow precipitation that were calibrated to support conservation of the water balance. The LSTM model, on the other hand, has no water balance constraints and has therefore no need for precipitation correction factors. We suggest that the improved bias ratios obtained by LSTM result from the absence of water balance constraints, allowing for a more accurate representation of average streamflow. Further, LSTM outperformed HBV in all but three catchments in terms of mean absolute error, implying a systematically higher accuracy in the LSTM predictions.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Within-model differences in flood peak performance for snowmelt versus rainfall generated floods</title>
      <p id="d2e3916">The results revealed notable differences in the ability of LSTM to simulate the correct peak timing of snowmelt versus rainfall generated floods. Whereas 77 % of rainfall generated flood peaks were simulated at the correct day, the corresponding percentage for snowmelt generated flood peaks was only 50 % (Fig. <xref ref-type="fig" rid="F5"/>). This finding was robust across the three different pairs of classification thresholds tested (Figs. <xref ref-type="fig" rid="FC3"/>–<xref ref-type="fig" rid="FC5"/>). As snowmelt generated floods react to snowmelt typically spanning days to weeks, the floods often last multiple days and may not have a very distinct peak as compared to rainfall generated floods <xref ref-type="bibr" rid="bib1.bibx6" id="paren.64"><named-content content-type="pre">e.g. Fig. 2 in</named-content></xref>. Relatively small differences between peak discharge and discharge magnitudes in adjacent days may explain the lower percentage of correctly simulated peak days for snowmelt generated floods. This reasoning is supported by the similar mean absolute percentage errors at the day of observed peak (<inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>) and FAR and POD results for snowmelt and rainfall generated floods (Fig. <xref ref-type="fig" rid="F8"/>). The same reasoning may explain the intermediate position of the timing results for mixed flood (better than snowmelt generated floods and worse than rainfall generated floods).</p>
      <p id="d2e3944">The identified differences in model performance for different flood types were not unique to LSTM. Similar differences were found for HBV. In terms of <inline-formula><mml:math id="M168" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, HBV exhibited the highest errors (exceeding 40 %) for rainfall generated floods in several catchments. Similarly high errors were only found in three catchments for LSTM. Flood-type dependent performance has also previously been demonstrated by <xref ref-type="bibr" rid="bib1.bibx10" id="text.65"/>, who identified differences in a model's flood peak performances in catchments of different regimes across four different process-based hydrological models.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Model evaluations reflecting the characteristics of interest</title>
      <p id="d2e3969">The value of a hydrological model for specific applications depends on the model performance with regards to relevant characteristics. Thus, model evaluations beyond the metrics considering the entire streamflow time series (e.g. NSE and KGE) are often necessary. Figure <xref ref-type="fig" rid="F10"/> shows that the model preferred for simulating the full streamflow time series (KGE) does not always match the one preferred for reproducing peak magnitudes (<inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>). Whereas the majority of the catchments had better LSTM scores for both overall and flood peak metrics (quadrant IV), the preferred model depended on the metric for approx. one fifth of the catchments (quadrants I and III). Notably, some catchments had a slightly (<inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">0.025</mml:mn></mml:mrow></mml:math></inline-formula>) lower KGE for LSTM than HBV, whereas the peak magnitude errors were reduced by 5 to 10 pp. In such cases, LSTM may be the preferred model for flood-specific applications despite the somewhat lower overall score.</p>
      <p id="d2e3995">We included three different flood peak specific evaluation metrics in this study to account for different aspects relevant for flood applications. Specifically, <inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> was evaluated both for simulated discharge at the day of observed peak (“pinned”), and for simulated maximum discharge within 2 d of the observed peak (“floating”). For HBV, which generally had higher timing errors, the differences between the two <inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> metrics were larger than for LSTM. However, the overweight of catchments with a better LSTM <inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> score make evident that peak magnitudes were better represented by LSTM regardless of the timing errors by HBV. Generally, our flood peak evaluation results demonstrated that LSTM does a better job than HBV for most catchments. At the local (i.e. catchment) scale, however, the conclusion on the preferred model depends in some cases on the desired characteristic, and the desired characteristics depend on the application. For issuing flood warning, for example, the timing may be crucial, whereas inundation mapping depend more on magnitudes being correctly simulated.</p>
      <p id="d2e4032">A related interesting model evaluation step, outside of scope of this study, is to evaluate to what degree flood event classification using data from the hydrological models correspond to the applied classification. However, such an evaluation has to be done differently for the two models. For HBV, flood classification using model simulated snowmelt and derived rainfall can be compared with the classification performed using SeNorge derived rainfall and snowmelt. LSTM, on the other hand, does not simulate snowmelt, and an evaluation of its ability to classify flood events into dominant FGP, would rely on explainable machine learning techniques. One such technique proposed by <xref ref-type="bibr" rid="bib1.bibx24" id="text.66"/> uses integrated gradients to trace back contributions of each input sequence (precipitation and temperature) to individual flood events, and applies cluster analysis to identify dominant flooding mechanisms. They identified three main FGPs across Europe, two of them dominating in Norway and strongly reflecting snowmelt and rainfall generated flood events, respectively. A primary aim of a comparison of such a classification approach with the one used here, could be to assess the robustness of classification technique and subjective choices regarding FGP definitions. The relatively simple classification technique and related subjective choices used in this study is subject to uncertainties, and a comparison with other techniques would be an interesting further study. That said, we expect limited effect of chosen technique on our main results and conclusions, as indicated by the very similar results produced using alternative classification thresholds.</p>
</sec>
<sec id="Ch1.S4.SS4">
  <label>4.4</label><title>Training LSTM on the premises of deep learning models</title>
      <p id="d2e4047">In this study, we constrained the dynamical forcing data and training period for LSTM to match that of our benchmark model in order to have a reasonably fair comparison. By doing so, we ensured that differences in model performances could not be attributed to differences in how informed the models were about local hydrometeorological conditions in each catchment. However, given the fundamentally different nature of the two models, the premises for training the best model are different. Accordingly, premade choices suitable of the benchmark model is not the best suitable choices for deep learning models, and there is a further potential to improve the LSTM simulations. Low-hanging fruit to achieve an LSTM model with even higher performance include expanding the training period considerably where possible and forcing the model with a wider range of atmospheric variables and datasets <xref ref-type="bibr" rid="bib1.bibx39 bib1.bibx32" id="paren.67"/>. Another possibility is finetuning of the LSTM model for individual catchments to explore the potential to improve the model further for local conditions <xref ref-type="bibr" rid="bib1.bibx29" id="paren.68"/>. Such potential avenues should be explored in case an LSTM model is considered for operational use to unleash the best potential based on the premises of deep learning models.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d2e4067">This study evaluated the ability of LSTM to simulate floods of different flood generating processes by assessing peak timing and peak magnitude of snowmelt generated floods and rainfall generated floods separately. To evaluate LSTM's potential for operational use in snow-influenced regions, the results were compared with the operational model in the study region, HBV. Our findings can be summarised by the following answers to our research questions:</p>
      <p id="d2e4070"><list list-type="order">
          <list-item>

      <p id="d2e4075">LSTM simulated streamflow series with higher performance than HBV for 96 % of the catchments in terms of NSE and 77 % of the catchments in terms of KGE. The largest improvements by LSTM were found for catchments with the lowest HBV performance scores.</p>
          </list-item>
          <list-item>

      <p id="d2e4081">LSTM simulated the correct timing of the flood peaks more often for rainfall generated events (77 %) than mixed events (67 %) and snowmelt generated events (50 %). The percentages of correctly simulated peak timing were 14–16 pp higher than those of HBV for all snowmelt, mixed and rainfall generated events, mainly because HBV more frequently simulated peak discharge too early.</p>

      <p id="d2e4084">LSTM simulated flood peak magnitudes with a similar performance for snowmelt generated floods, mixed floods and rainfall generated floods. Percent errors were within <inline-formula><mml:math id="M174" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % for close to half of the events, a slightly better result than that of HBV. Both LSTM and HBV generally underestimated flood peak magnitudes, whereas HBV had a higher proportion of the largest underestimations (more than 40 % underestimation) as compared to LSTM for the three types of floods. Resulting false alarm rates and probabilities of detection were similar for rainfall and snowmelt generated events, and slightly better than those of mixed flood events.</p>
          </list-item>
          <list-item>

      <p id="d2e4100">A large spread among catchments was found in timing errors by LSTM for snowmelt generated floods (approx. 10 % to 80 %), with an average of 49 %. Timing errors for rainfall generated floods were notably smaller, with an average error of 24 % and only two catchments exceeding 50 %. LSTM exhibited smaller timing errors than HBV in the majority of the catchments for both snowmelt and rainfall generated floods.</p>

      <p id="d2e4104">Per-catchment LSTM results of mean absolute percentage error of flood peaks were mainly in the range 10 % to 40 % for both snowmelt and rainfall generated flood events. For both types of floods, the errors were smaller for LSTM as compared to HBV for the majority of the catchments. Most model differences were within 16 pp, but LSTM improved the rainfall generated flood peak magnitude simulations with up to 33 pp for catchments with relatively poor HBV scores.</p>
          </list-item>
        </list></p>
      <p id="d2e4109">Overall, LSTM provided reliable simulations of streamflow time series and floods of different flood generating processes in catchments influenced by seasonal snow. Notable improvements were found in the timing of simulated flood peaks, and in overall and peak magnitude metrics for catchments where the benchmark model had relatively poor results. Our findings highlight LSTM's potential to improve hydrological services and flood assessments in regions prone to both snowmelt and rainfall generated floods.</p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <label>Appendix A</label><title>Training the LSTM model</title>
      <p id="d2e4123">The LSTM model was trained using the 15-years period 1 September 2009–31 August 2024, to match the calibration period used for our benchmark model simulations. The catchment-averaged NSE, used as loss function, was minimised using the Adam optimiser <xref ref-type="bibr" rid="bib1.bibx27" id="paren.69"/> and we used a linear output activation function. The forget gate bias was initialised to 3, a common practice that encourages the LSTM to retain information early in training <xref ref-type="bibr" rid="bib1.bibx15 bib1.bibx16" id="paren.70"/>. To stabilise training, gradient norms were clipped to a maximum value of 1. The model was validated at every epoch, evaluating performance on 50 randomly selected catchments from the validation set. To mitigate overfitting during training, the final hidden state vector <inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:msub><mml:mi>h</mml:mi><mml:mi mathvariant="normal">t</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> was passed through a dropout layer, which randomly sets each element to zero with probability <inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:msub><mml:mi>p</mml:mi><mml:mtext>dropout</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e4154">The hidden state dimension, dropout probability, batch size (the number of training examples used to compute the loss gradient in each model update), input sequence length and learning rate were treated as tunable hyperparameters. These hyperparameters were tuned in a 5-fold cross-validation (CV) over the period 1 September 2009–31 August 2024. For each of the five CV-iterations, the period was split into a validation period of three consecutive years, and a left-out-fold of three consecutive years. Each year was used as left-out-fold once and validation period once during the CV. The years adjacent to the left-out-fold were removed to have a temporal “buffer” between the left-out-fold and the “seen” years, and the remaining years comprised the training period in each CV-iteration. The following hyperparameter values were considered in a grid search in each CV split (selected values for final model in bold):</p>
      <p id="d2e4157"><list list-type="bullet">
          <list-item>

      <p id="d2e4162">Hidden state dimension <inline-formula><mml:math id="M177" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> [<bold>64</bold>, 128, 256]</p>
          </list-item>
          <list-item>

      <p id="d2e4178">Dropout probability <inline-formula><mml:math id="M178" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> [0.2, 0.3, <bold>0.4</bold>]</p>
          </list-item>
          <list-item>

      <p id="d2e4194">Batch size <inline-formula><mml:math id="M179" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> [64, 128, <bold>256</bold>]</p>
          </list-item>
          <list-item>

      <p id="d2e4210">Input sequence length <inline-formula><mml:math id="M180" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> [<bold>270</bold>, 365]</p>
          </list-item>
          <list-item>

      <p id="d2e4226">Learning rate <inline-formula><mml:math id="M181" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> [<bold>0: 1e-3  10: 5e-4  25: 1e-4</bold>, 0: 5e-3  10: 1e-3  25: 5e-4]</p>
          </list-item>
        </list></p>
      <p id="d2e4241">where the learning rate schedule notation indicates the learning rate value at specified epochs (i.e. epochs 0, 10, and 25). This yielded 108 hyperparameter combinations per cross-validation split. The number of training epochs (from 1 to 50) was selected by evaluating validation NSE at each epoch for each hyperparameter combination.</p>
      <p id="d2e4245">Models using the faster learning rate schedule frequently exhibited poor performance (low or negative validation NSE) and training instability, leading us to exclude these configurations from further consideration. For the remaining 54 hyperparameter combinations per split, we evaluated performance on the left-out-fold using the three epochs with highest catchment-averaged validation NSE: epoch 35 (NSE <inline-formula><mml:math id="M182" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.852), epoch 26 (NSE <inline-formula><mml:math id="M183" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.850), and epoch 10 (NSE <inline-formula><mml:math id="M184" display="inline"><mml:mo>=</mml:mo></mml:math></inline-formula> 0.849).</p>
      <p id="d2e4269">The final hyperparameter configuration (shown in bold above) was selected as the combination yielding the highest average NSE across all left-out-folds and evaluated epochs. The final model was trained using 1 September 2009–31 August 2019 as the training period and 1 September 2019–31 August 2024 as the validation period. We selected epoch 35 for the final model as it achieved the highest catchment-averaged validation NSE of 0.87.</p>
</app>

<app id="App1.Ch1.S2">
  <label>Appendix B</label><title>HBV calibration</title>
      <p id="d2e4280">For descriptions of the version and set-up of the daily HBV model used operationally in Norway, we refer to <xref ref-type="bibr" rid="bib1.bibx35" id="text.71"/> and <xref ref-type="bibr" rid="bib1.bibx46" id="text.72"/>, except that the model has been re-calibrated for a more recent period (1 September 2009–31 August 2024) using a more recent meteorological dataset (SeNorge_2018) after the two reports were published. In short, the applied HBV model is a semi-distributed bucket-type model, that divides the catchment into ten elevation zones to account for elevation gradients in temperature and precipitation <xref ref-type="bibr" rid="bib1.bibx48 bib1.bibx26" id="paren.73"/>. The model uses elevation zone averaged daily precipitation sums and mean temperatures as input. Fluxes between and processes within four storage components, i.e. snow, soil moisture, an upper runoff zone, and a lower runoff zone, are represented by simplified process-based expressions.</p>
      <p id="d2e4292">We calibrated the HBV model using the operational model's calibration set-up and choices, except that we used Nash-Sutcliffe efficiency (NSE) as a loss function to match that of LSTM. Catchment-by-catchment calibration was conducted using Model-Independent Parameter Estimation and Uncertainty Analysis <xref ref-type="bibr" rid="bib1.bibx11" id="paren.74"><named-content content-type="pre">PEST;</named-content></xref> with parameter ranges from Table 1 in <xref ref-type="bibr" rid="bib1.bibx46" id="text.75"/>. Daily discharge and meteorological observations in Norway have historically been recorded at two different daily divides, i.e. 00:00–00:00 CET and 07:00-07:00 CET, respectively. In the HBV model used operationally, discharge and meteorological data are matched using the largest overlap possible (17 h). Correspondingly, we kept this choice for the modelling of both HBV and LSTM.</p>
</app>

<app id="App1.Ch1.S3">
  <label>Appendix C</label><title>Additional figures</title>

      <fig id="FC1"><label>Figure C1</label><caption><p id="d2e4313">Runoff coefficient (i.e. ratio of streamflow (mm d<sup>−1</sup>) to precipitation (mm d<sup>−1</sup>)) of each catchment in the training period (i.e. calibration period; <inline-formula><mml:math id="M187" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) versus the evaluation period (<inline-formula><mml:math id="M188" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis). A dot close to the diagonal line implies consistent runoff coefficients for the two periods. Catchments with glaciers covering more than 3 % of their area are marked with stars. We note that values exceeding 1 (i.e. streamflow larger than precipitation) are mainly due to underestimation of precipitation, although glacier melt can explain part of the difference in catchments with glaciers. Each catchment is coloured by the corresponding catchment's dominant flood generating process (FGP, ref. Fig. <xref ref-type="fig" rid="F2"/>).</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f12.png"/>

      </fig>

<fig id="FC2"><label>Figure C2</label><caption><p id="d2e4367">Overall performance per catchment for HBV (<inline-formula><mml:math id="M189" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) versus LSTM (<inline-formula><mml:math id="M190" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis) in terms of the three KGE components <bold>(a)</bold> correlation coefficient, <inline-formula><mml:math id="M191" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula>, <bold>(b)</bold> bias ratio, <inline-formula><mml:math id="M192" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula>, and <bold>(c)</bold> variability ratio, <inline-formula><mml:math id="M193" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>, as well as <bold>(d)</bold> mean absolute error, MAE. Optimum values are marked with stippled lines. Backgrounds are shaded by the corresponding 2D empirical density functions. Triangles with adjacent numbers represent catchment average score for each model. The percentages of catchments above and below the diagonal lines are provided in the upper right corner of the scatterplots of <inline-formula><mml:math id="M194" display="inline"><mml:mi mathvariant="italic">ρ</mml:mi></mml:math></inline-formula> and MAE. For <inline-formula><mml:math id="M195" display="inline"><mml:mi mathvariant="italic">β</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M196" display="inline"><mml:mi mathvariant="italic">γ</mml:mi></mml:math></inline-formula>, percentages of scores below and above optimum value of 1 are provided for both models. Each catchment is coloured by the corresponding catchment's dominant flood generating process (FGP, ref. Fig. <xref ref-type="fig" rid="F2"/>).</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f13.png"/>

      </fig>

<fig id="FC3"><label>Figure C3</label><caption><p id="d2e4454">Distribution of all flood events' fractional rainfall contribution. Dark green vertical lines represent flood type classification thresholds used for the main analyses. Light green and blue vertical lines represent two alternative pairs of classification thresholds for which Figs. <xref ref-type="fig" rid="F5"/> (flood peak timing) and <xref ref-type="fig" rid="F6"/> (flood peak magnitude) have been repeated (see Figs. <xref ref-type="fig" rid="FC4"/>–<xref ref-type="fig" rid="FC7"/>).</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f14.png"/>

      </fig>

      <fig id="FC4"><label>Figure C4</label><caption><p id="d2e4475">As Fig. <xref ref-type="fig" rid="F5"/>, but using the alternative symmetric thresholds 0.2 and 0.8 for flood type classification (see Fig. <xref ref-type="fig" rid="FC3"/>). Note that the values in parenthesis, representing the number of flood events, are different from Fig. <xref ref-type="fig" rid="F5"/>.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f15.png"/>

      </fig>

<fig id="FC5"><label>Figure C5</label><caption><p id="d2e4495">As Fig. <xref ref-type="fig" rid="F5"/>, but using the alternative asymmetric thresholds 0.3 and 0.9 for flood type classification (see Fig. <xref ref-type="fig" rid="FC3"/>). Note that the values in parenthesis, representing the number of flood events, are different from Fig. <xref ref-type="fig" rid="F5"/>.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f16.png"/>

      </fig>

      <fig id="FC6"><label>Figure C6</label><caption><p id="d2e4514">As Fig. <xref ref-type="fig" rid="F6"/>, but using the alternative symmetric thresholds 0.2 and 0.8 for flood type classification (see Fig. <xref ref-type="fig" rid="FC3"/>).</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f17.png"/>

      </fig>

<fig id="FC7"><label>Figure C7</label><caption><p id="d2e4533">As Fig. <xref ref-type="fig" rid="F6"/>, but using the alternative asymmetric thresholds 0.3 and 0.9 for flood type classification (see Fig. <xref ref-type="fig" rid="FC3"/>).</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f18.png"/>

      </fig>

<fig id="FC8"><label>Figure C8</label><caption><p id="d2e4551">Model difference (i.e. LSTM minus HBV) in the overall performance metric NSE (<inline-formula><mml:math id="M197" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) versus the flood peak metrics: <bold>(a–b)</bold> timing error, <bold>(c–d)</bold> <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <bold>(e–f)</bold> <inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, separately for snowmelt generated floods (left) and rainfall generated floods (right). Each dot represents a catchment and is coloured by the catchment's dominant flood generating process (FGP; ref. Fig. <xref ref-type="fig" rid="F2"/>). Catchments within quadrant IV (QIV; lower right square) have a better score for LSTM as compared to HBV both in terms of overall performance score (NSE) and flood peak performance score. Percentages of catchments within each quadrant are given in the corners of the plots.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f19.png"/>

      </fig>

<fig id="FC9"><label>Figure C9</label><caption><p id="d2e4606">Model difference (i.e. LSTM minus HBV) in the overall performance metric KGE (<inline-formula><mml:math id="M200" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) versus the flood peak metrics: <bold>(a–b)</bold> timing error, <bold>(c–d)</bold> <inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,pinned</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <bold>(e–f)</bold> <inline-formula><mml:math id="M202" display="inline"><mml:mrow><mml:msub><mml:mtext>MAPE</mml:mtext><mml:mtext>peak,floating</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>, separately for snowmelt generated floods (left) and rainfall generated floods (right). Each dot represents a catchment and is coloured by the catchment's dominant flood generating process (FGP; ref. Fig. <xref ref-type="fig" rid="F2"/>). Catchments within quadrant IV (QIV; lower right square) have a better score for LSTM as compared to HBV both in terms of overall performance score (KGE) and flood peak performance score. Percentages of catchments within each quadrant are given in the corners of the plots.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/6207/2026/hess-30-6207-2026-f20.png"/>

      </fig>


</app>
  </app-group><notes notes-type="dataavailability"><title>Data availability</title>

      <p id="d2e4664">Discharge data are openly available at <uri>https://hydapi.nve.no/</uri> (last access: 10 February 2026) <xref ref-type="bibr" rid="bib1.bibx42" id="paren.76"/>. SeNorge_2018 data are openly available at <uri>https://thredds.met.no/thredds/catalog/senorge/seNorge_2018/catalog.html</uri> (last access: 11 February 2026) <xref ref-type="bibr" rid="bib1.bibx40" id="paren.77"/>. The snowmelt dataset is openly available at <uri>https://thredds.met.no/thredds/catalog/senorge/seNorge_snow/qsw/catalog.html</uri> (last access: 10 February 2026) <xref ref-type="bibr" rid="bib1.bibx44" id="paren.78"/>. Prepared catchment-level data used in this study are made available at <ext-link xlink:href="https://doi.org/10.5281/zenodo.22942372" ext-link-type="DOI">10.5281/zenodo.22942372</ext-link> <xref ref-type="bibr" rid="bib1.bibx4" id="paren.79"/>.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e4695">SJB, DMB and SN designed the study with contributions from SAK and KE. All authors collected the data. KE quality controlled the catchment attributes. DMB calibrated HBV. SJB carried out the data preprocessing, LSTM modelling, analyses and visualisations with contributions from DMB. SJB, with input from DMB, wrote the original draft, and all authors contributed to revision and editing of the manuscript.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e4701">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e4707">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e4713">Data and code providers are greatly acknowledged. We thank the Norwegian Water Resources and Energy Directorate (NVE) for providing gridded snowmelt data, catchment attributes, observed streamflow data, and the model set-up of the operational HBV model. We also thank the Norwegian Meteorological institute for providing the seNorge_2018 data. We further thank the team behind neuralhydrology, who make LSTM modelling accessible for the broader hydrological community.</p></ack><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e4718">This paper was edited by Thom Bogaard and reviewed by Klaus Vormoor and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Addor and Melsen(2019)</label><mixed-citation>Addor, N. and Melsen, L. A.: Legacy, Rather Than Adequacy, Drives the Selection of Hydrological Models, Water Resour. Res., 55, 378–390, <ext-link xlink:href="https://doi.org/10.1029/2018WR022958" ext-link-type="DOI">10.1029/2018WR022958</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Anderson and Radić(2022)</label><mixed-citation>Anderson, S. and Radić, V.: Evaluation and interpretation of convolutional long short-term memory networks for regional hydrological modelling, Hydrol. Earth Syst. Sci., 26, 795–825, <ext-link xlink:href="https://doi.org/10.5194/hess-26-795-2022" ext-link-type="DOI">10.5194/hess-26-795-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Andreassen et al.(2022)Andreassen, Nagy, Kjøllmoen, and Leigh</label><mixed-citation>Andreassen, L. M., Nagy, T., Kjøllmoen, B., and Leigh, J. R.: An inventory of Norway's glaciers and ice-marginal lakes from 2018–19 Sentinel-2 data, J. Glaciol., 68, 1085–1106, <ext-link xlink:href="https://doi.org/10.1017/jog.2022.20" ext-link-type="DOI">10.1017/jog.2022.20</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Bakke et al.(2026)</label><mixed-citation>Bakke, S. J., Barna, D. M., Engeland, K., Kolberg, S. A., and Nordeide, S.: Data for “The ability of LSTM to model snowmelt versus rainfall generated floods”, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.22942372" ext-link-type="DOI">10.5281/zenodo.22942372</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Barna et al.(2023a)Barna, Engeland, Kneib, Thorarinsdottir, and Xu</label><mixed-citation>Barna, D. M., Engeland, K., Kneib, T., Thorarinsdottir, T. L., and Xu, C.-Y.: Regional index flood estimation at multiple durations with generalized additive models, EGUsphere [preprint], <ext-link xlink:href="https://doi.org/10.5194/egusphere-2023-2335" ext-link-type="DOI">10.5194/egusphere-2023-2335</ext-link>, 2023a.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Barna et al.(2023b)Barna, Engeland, Thorarinsdottir, and Xu</label><mixed-citation>Barna, D. M., Engeland, K., Thorarinsdottir, T. L., and Xu, C.-Y.: Flexible and consistent Flood–Duration–Frequency modeling: A Bayesian approach, J. Hydrol., 620, 129448, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2023.129448" ext-link-type="DOI">10.1016/j.jhydrol.2023.129448</ext-link>, 2023b.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Berghuijs et al.(2019)Berghuijs, Harrigan, Molnar, Slater, and Kirchner</label><mixed-citation>Berghuijs, W. R., Harrigan, S., Molnar, P., Slater, L. J., and Kirchner, J. W.: The Relative Importance of Different Flood-Generating Mechanisms Across Europe, Water Resour. Res., 55, 4582–4593, <ext-link xlink:href="https://doi.org/10.1029/2019WR024841" ext-link-type="DOI">10.1029/2019WR024841</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Bergström(1976)</label><mixed-citation> Bergström, S.: Development and application of a conceptual runoff model for Scandinavian catchments, Tech. rep., SMHI Report Nr RHO 7/1976, The Swedish Meteorological and Hydrological Institute (SMHI), ISSN 0347-7827, 1976.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Bocharov(2023)</label><mixed-citation>Bocharov, G.: pyextremes, <uri>https://github.com/georgebv/pyextremes</uri> (last access: 2 December 2025), 2023.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Brunner et al.(2021)Brunner, Melsen, Wood, Rakovec, Mizukami, Knoben, and Clark</label><mixed-citation>Brunner, M. I., Melsen, L. A., Wood, A. W., Rakovec, O., Mizukami, N., Knoben, W. J. M., and Clark, M. P.: Flood spatial coherence, triggers, and performance in hydrological simulations: large-sample evaluation of four streamflow-calibrated models, Hydrol. Earth Syst. Sci., 25, 105–119, <ext-link xlink:href="https://doi.org/10.5194/hess-25-105-2021" ext-link-type="DOI">10.5194/hess-25-105-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Doherty(2004)</label><mixed-citation>Doherty, J.: PEST: Model Independent Parameter Estimation, User Manual, 5th edn., <uri>https://www.nrc.gov/docs/ML0923/ML092360221.pdf</uri> (last access: 2 October 2026), 2004.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Dyrrdal et al.(2025)Dyrrdal, Bakke, Hanssen-Bauer, Mayer, Nilsen, Nilsen, Paasche, Saloranta, and Årthun</label><mixed-citation>Dyrrdal, A., Bakke, S., Hanssen-Bauer, I., Mayer, S., Nilsen, I., Nilsen, J., Paasche, Ø., Saloranta, T., and Årthun, M. (Eds.): Klima i Norge – kunnskapsgrunnlag for klimatilpasning oppdatert i 2025, NCCS-rapport 1/2025, The Norwegian Center for Climate Services (NCCS), <ext-link xlink:href="https://doi.org/10.60839/4rgq-nn84" ext-link-type="DOI">10.60839/4rgq-nn84</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Engeland et al.(2016)Engeland, Schlichting, Randen, Nordtun, Reitan, Wang, Holmqvist, Voksø, and Eide</label><mixed-citation>Engeland, K., Schlichting, L., Randen, F., Nordtun, K., Reitan, T., Wang, T., Holmqvist, E., Voksø, A., and Eide, V.: Utvalg og kvalitetssikring av flomdata forflomfrekvensanalyser, Technical Report 85/2016, The Norwegian Water Resources and Energy Directorate (NVE), ISBN 978-82-410-1538-0, <uri>https://publikasjoner.nve.no/rapport/2016/rapport2016_85.pdf</uri> (last access: 2 October 2026), 2016.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Engeland et al.(2020)Engeland, Glad, Hamududu, Li, Reitan, and Stenius</label><mixed-citation>Engeland, K., Glad, P., Hamududu, B. H., Li, H., Reitan, T., and Stenius, S. M.: Lokal og regional flomfrekvensanalyse, Technical Report 10/2020, The Norwegian Water Resources and Energy Directorate (NVE), ISBN 978-82-410-2014-8, <uri>https://publikasjoner.nve.no/rapport/2020/rapport2020_10.pdf</uri> (last access: 2 October 2026), 2020.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Frame et al.(2022)Frame, Kratzert, Klotz, Gauch, Shalev, Gilon, Qualls, Gupta, and Nearing</label><mixed-citation>Frame, J. M., Kratzert, F., Klotz, D., Gauch, M., Shalev, G., Gilon, O., Qualls, L. M., Gupta, H. V., and Nearing, G. S.: Deep learning rainfall–runoff predictions of extreme events, Hydrol. Earth Syst. Sci., 26, 3377–3392, <ext-link xlink:href="https://doi.org/10.5194/hess-26-3377-2022" ext-link-type="DOI">10.5194/hess-26-3377-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Gauch et al.(2021)Gauch, Kratzert, Klotz, Nearing, Lin, and Hochreiter</label><mixed-citation>Gauch, M., Kratzert, F., Klotz, D., Nearing, G., Lin, J., and Hochreiter, S.: Rainfall–runoff prediction at multiple timescales with a single Long Short-Term Memory network, Hydrol. Earth Syst. Sci., 25, 2045–2062, <ext-link xlink:href="https://doi.org/10.5194/hess-25-2045-2021" ext-link-type="DOI">10.5194/hess-25-2045-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>GeoNorge(2026a)</label><mixed-citation>GeoNorge: Totalnedbørfelt til målestasjon, The Norwegian Mapping Authority (Kartverket), <uri>https://kartkatalog.geonorge.no/metadata/totalnedboerfelt-til-maalestasjon/ac1c71db-9850-4e89-8162-2baba8b980e7</uri> (last access: 10 February 2026), 2026a.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>GeoNorge(2026b)</label><mixed-citation>GeoNorge: ELVIS elvenett, The Norwegian Mapping Authority (Kartverket), <uri>https://kartkatalog.geonorge.no/metadata/elvis-elvenett/3f95a194-0968-4457-a500-912958de3d39</uri> (last access: 10 February 2026), 2026b.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>GeoNorge(2026c)</label><mixed-citation>GeoNorge: Løsmasser, infiltrasjonEvne, The Norwegian Mapping Authority (Kartverket), <uri>https://kartkatalog.geonorge.no/metadata/loesmasser/3de4ddf6-d6b8-4398-8222-f5c47791a757</uri>, (last access: 10 February 2026), 2026c.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Gers et al.(2000)Gers, Schmidhuber, and Cummins</label><mixed-citation>Gers, F. A., Schmidhuber, J., and Cummins, F.: Learning to Forget: Continual Prediction with LSTM, Neural Comput., 12, 2451–2471, <ext-link xlink:href="https://doi.org/10.1162/089976600300015015" ext-link-type="DOI">10.1162/089976600300015015</ext-link>, 2000.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Gupta et al.(2009)Gupta, Kling, Yilmaz, and Martinez</label><mixed-citation>Gupta, H. V., Kling, H., Yilmaz, K. K., and Martinez, G. F.: Decomposition of the mean squared error and NSE performance criteria: Implications for improving hydrological modelling, J. Hydrol., 377, 80–91, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2009.08.003" ext-link-type="DOI">10.1016/j.jhydrol.2009.08.003</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Hagen et al.(2023)Hagen, Hasibi, Leblois, Lawrence, and Sorteberg</label><mixed-citation>Hagen, J. S., Hasibi, R., Leblois, E., Lawrence, D., and Sorteberg, A.: Reconstructing daily streamflow and floods from large-scale atmospheric variables with feed-forward and recurrent neural networks in high latitude climates, Hydrolog. Sci. J., 68, 412–431, <ext-link xlink:href="https://doi.org/10.1080/02626667.2023.2165927" ext-link-type="DOI">10.1080/02626667.2023.2165927</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Hochreiter and Schmidhuber(1997)</label><mixed-citation>Hochreiter, S. and Schmidhuber, J.: Long short-term memory, Neural Compu., 9, 1735–1780, <ext-link xlink:href="https://doi.org/10.1162/neco.1997.9.8.1735" ext-link-type="DOI">10.1162/neco.1997.9.8.1735</ext-link>, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Jiang et al.(2022)Jiang, Bevacqua, and Zscheischler</label><mixed-citation>Jiang, S., Bevacqua, E., and Zscheischler, J.: River flooding mechanisms and their changes in Europe revealed by explainable machine learning, Hydrol. Earth Syst. Sci., 26, 6339–6359, <ext-link xlink:href="https://doi.org/10.5194/hess-26-6339-2022" ext-link-type="DOI">10.5194/hess-26-6339-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Kartverket(2026)</label><mixed-citation>Kartverket: Høydedata, The Norwegian Mapping Authority (Kartverket), <uri>https://hoydedata.no/LaserInnsyn2/</uri> (last access: 10 February 2026), 2026.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Killingtveit and Saelthun(1995)</label><mixed-citation> Killingtveit, Å and Saelthun, N. R.: Hydrological models, Hydropower development, Norwegian Inst. of Technology, Dept. of Hydraulic Engineering, Vol. 7, 99–128, ISBN 978-82-7598-026-5, 1995.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Kingma and Ba(2017)</label><mixed-citation>Kingma, D. P. and Ba, J.: Adam: A Method for Stochastic Optimization. Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015), San Diego, CA, USA, <uri>https://mlanthology.org/iclr/2015/kingma2015iclr-adam/</uri> (2 October 2026), 2015.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Kling et al.(2012)Kling, Fuchs, and Paulin</label><mixed-citation>Kling, H., Fuchs, M., and Paulin, M.: Runoff conditions in the upper Danube basin under an ensemble of climate change scenarios, J. Hydrol., 424—425, 264–277, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2012.01.011" ext-link-type="DOI">10.1016/j.jhydrol.2012.01.011</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Kratzert et al.(2018)Kratzert, Klotz, Brenner, Schulz, and Herrnegger</label><mixed-citation>Kratzert, F., Klotz, D., Brenner, C., Schulz, K., and Herrnegger, M.: Rainfall–runoff modelling using Long Short-Term Memory (LSTM) networks, Hydrol. Earth Syst. Sci., 22, 6005–6022, <ext-link xlink:href="https://doi.org/10.5194/hess-22-6005-2018" ext-link-type="DOI">10.5194/hess-22-6005-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Kratzert et al.(2019a)Kratzert, Herrnegger, Klotz, Hochreiter, and Klambauer</label><mixed-citation>Kratzert, F., Herrnegger, M., Klotz, D., Hochreiter, S., and Klambauer, G.: NeuralHydrology – Interpreting LSTMs in Hydrology, Springer International Publishing, Cham, 347–362,, ISBN 978-3-030-28954-6, <ext-link xlink:href="https://doi.org/10.1007/978-3-030-28954-6_19" ext-link-type="DOI">10.1007/978-3-030-28954-6_19</ext-link>, 2019a.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Kratzert et al.(2019b)Kratzert, Klotz, Shalev, Klambauer, Hochreiter, and Nearing</label><mixed-citation>Kratzert, F., Klotz, D., Shalev, G., Klambauer, G., Hochreiter, S., and Nearing, G.: Towards learning universal, regional, and local hydrological behaviors via machine learning applied to large-sample datasets, Hydrol. Earth Syst. Sci., 23, 5089–5110, <ext-link xlink:href="https://doi.org/10.5194/hess-23-5089-2019" ext-link-type="DOI">10.5194/hess-23-5089-2019</ext-link>, 2019b.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Kratzert et al.(2021)Kratzert, Klotz, Hochreiter, and Nearing</label><mixed-citation>Kratzert, F., Klotz, D., Hochreiter, S., and Nearing, G. S.: A note on leveraging synergy in multiple meteorological data sets with deep learning for rainfall–runoff modeling, Hydrol. Earth Syst. Sci., 25, 2685–2703, <ext-link xlink:href="https://doi.org/10.5194/hess-25-2685-2021" ext-link-type="DOI">10.5194/hess-25-2685-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Kratzert et al.(2022)Kratzert, Gauch, Nearing, and Klotz</label><mixed-citation>Kratzert, F., Gauch, M., Nearing, G., and Klotz, D.: NeuralHydrology – A Python library for Deep Learning research in hydrology, Journal of Open Source Software, 7, 4050, <ext-link xlink:href="https://doi.org/10.21105/joss.04050" ext-link-type="DOI">10.21105/joss.04050</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Kratzert et al.(2024)Kratzert, Gauch, Klotz, and Nearing</label><mixed-citation>Kratzert, F., Gauch, M., Klotz, D., and Nearing, G.: HESS Opinions: Never train a Long Short-Term Memory (LSTM) network on a single basin, Hydrol. Earth Syst. Sci., 28, 4187–4201, <ext-link xlink:href="https://doi.org/10.5194/hess-28-4187-2024" ext-link-type="DOI">10.5194/hess-28-4187-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Lawrence et al.(2009)Lawrence, Haddeland, and Langsholt</label><mixed-citation> Lawrence, D., Haddeland, I., and Langsholt, E.: Calibration of HBV hydrological models using PEST parameter estimation, Technical Report 1/2009, The Norwegian Water Resources and Energy Directorate (NVE), ISBN 78-82-410-0680-7, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Lees et al.(2022)Lees, Reece, Kratzert, Klotz, Gauch, De Bruijn, Kumar Sahu, Greve, Slater, and Dadson</label><mixed-citation>Lees, T., Reece, S., Kratzert, F., Klotz, D., Gauch, M., De Bruijn, J., Kumar Sahu, R., Greve, P., Slater, L., and Dadson, S. J.: Hydrological concept formation inside long short-term memory (LSTM) networks, Hydrol. Earth Syst. Sci., 26, 3079–3101, <ext-link xlink:href="https://doi.org/10.5194/hess-26-3079-2022" ext-link-type="DOI">10.5194/hess-26-3079-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Liu et al.(2025)Liu, Shen, O'Donncha, Song, Zhi, Beck, Bindas, Kraabel, and Lawson</label><mixed-citation>Liu, J., Shen, C., O'Donncha, F., Song, Y., Zhi, W., Beck, H. E., Bindas, T., Kraabel, N., and Lawson, K.: From RNNs to Transformers: benchmarking deep learning architectures for hydrologic prediction, Hydrol. Earth Syst. Sci., 29, 6811–6828, <ext-link xlink:href="https://doi.org/10.5194/hess-29-6811-2025" ext-link-type="DOI">10.5194/hess-29-6811-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Lussana et al.(2019)Lussana, Tveito, Dobler, and Tunheim</label><mixed-citation>Lussana, C., Tveito, O. E., Dobler, A., and Tunheim, K.: seNorge_2018, daily precipitation, and temperature datasets over Norway, Earth Syst. Sci. Data, 11, 1531–1551, <ext-link xlink:href="https://doi.org/10.5194/essd-11-1531-2019" ext-link-type="DOI">10.5194/essd-11-1531-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Martel et al.(2025)Martel, Arsenault, Turcotte, Castañeda Gonzalez, Brissette, Armstrong, Mailhot, Pelletier-Dumont, Lachance-Cloutier, Rondeau-Genesse, and Caron</label><mixed-citation>Martel, J.-L., Arsenault, R., Turcotte, R., Castañeda-Gonzalez, M., Brissette, F., Armstrong, W., Mailhot, E., Pelletier-Dumont, J., Lachance-Cloutier, S., Rondeau-Genesse, G., and Caron, L.-P.: Exploring the ability of LSTM-based hydrological models to simulate streamflow time series for flood frequency analysis, Hydrol. Earth Syst. Sci., 29, 4951–4968, <ext-link xlink:href="https://doi.org/10.5194/hess-29-4951-2025" ext-link-type="DOI">10.5194/hess-29-4951-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>MET Norway(2026)</label><mixed-citation>MET Norway: SeNorge_2018, The Norwegian Meteorological institute (MET Norway), <uri>https://thredds.met.no/thredds/catalog/senorge/seNorge_2018/catalog.html</uri> (last access: 11 February 2026), 2026.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Nearing et al.(2024)Nearing, Cohen, Dube, Gauch, Gilon, Harrigan, Hassidim, Klotz, Kratzert, Metzger et al.</label><mixed-citation>Nearing, G., Cohen, D., Dube, V., Gauch, M., Gilon, O., Harrigan, S., Hassidim, A., Klotz, D., Kratzert, F., Metzger, A., Nevo, S., Pappenberger, F., Prudhomme, C., Shalev, G., Shenzis, S., Tekalign, T. Y., Weitzner, D., and Matias, Y.: Global prediction of extreme floods in ungauged watersheds, Nature, 627, 559–563, <ext-link xlink:href="https://doi.org/10.1038/s41586-024-07145-1" ext-link-type="DOI">10.1038/s41586-024-07145-1</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>NVE(2026a)</label><mixed-citation>NVE: NVE Hydrological API (HydAPI), The Norwegian Water Resources and Energy Directorate (NVE), <uri>https://hydapi.nve.no/</uri> (last access: 10 February 2026), 2026a.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>NVE(2026b)</label><mixed-citation>NVE: Sildre, The Norwegian Water Resources and Energy Directorate (NVE), <uri>https://sildre.nve.no/</uri> (last access: 10 February 2026), 2026b.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>NVE and MET Norway(2026)</label><mixed-citation>NVE and MET Norway: Snowmelt, The Norwegian Water Resources and Energy Directorate (NVE) and the Norwegian Meteorological institute (MET Norway), <uri>https://thredds.met.no/thredds/catalog/senorge/seNorge_2018/catalog.html</uri> (last access: 10 February 2026), 2026.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Roksvåg et al.(2026)Roksvåg, Vandeskog, Wulff, and Wergeland</label><mixed-citation>Roksvåg, T., Vandeskog, S. M., Wulff, C., and Wergeland, K.: An LSTM network for joint modeling of streamflow and hydropower generation for run-of-river plants, J. Hydrol., 667, 134890, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2025.134890" ext-link-type="DOI">10.1016/j.jhydrol.2025.134890</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>Ruan and Langsholt(2017)</label><mixed-citation>Ruan, G. and Langsholt, E.: Rekalibrering av flomvarslingas HBV-modeller med inndata fra seNorge, versjon 2.0, Technical Report 71/2017, The Norwegian Water Resources and Energy Directorate (NVE), ISBN 978-82-410-1624-0, <uri>https://publikasjoner.nve.no/rapport/2017/rapport2017_71.pdf</uri> (last access: 2 October 2026), 2017.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Ruzzante et al.(2026)Ruzzante, Knoben, Wagener, Gleeson, and Schnorbus</label><mixed-citation>Ruzzante, S. W., Knoben, W. J. M., Wagener, T., Gleeson, T., and Schnorbus, M.: Technical note: High Nash–Sutcliffe Efficiencies conceal poor simulations of interannual variance in seasonal regimes, Hydrol. Earth Syst. Sci., 30, 2337–2355, <ext-link xlink:href="https://doi.org/10.5194/hess-30-2337-2026" ext-link-type="DOI">10.5194/hess-30-2337-2026</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx48"><label>Sælthun(1996)</label><mixed-citation>Sælthun, N. R.: The Nordic HBV Model, Technical Report 7/1996, The Norwegian Water Resources and Energy Directorate (NVE), ISBN 82-410-0273-4, <uri>https://publikasjoner.nve.no/publication/1996/publication1996_07.pdf</uri> (last access: 2 October 2026), 1996.</mixed-citation></ref>
      <ref id="bib1.bibx49"><label>Saloranta(2016)</label><mixed-citation>Saloranta, T. M.: Operational snow mapping with simplified data assimilation using the seNorge snow model, J. Hydrol., 538, 314–325, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2016.03.061" ext-link-type="DOI">10.1016/j.jhydrol.2016.03.061</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx50"><label>Seibert and Bergström(2022)</label><mixed-citation>Seibert, J. and Bergström, S.: A retrospective on hydrological catchment modelling based on half a century with the HBV model, Hydrol. Earth Syst. Sci., 26, 1371–1388, <ext-link xlink:href="https://doi.org/10.5194/hess-26-1371-2022" ext-link-type="DOI">10.5194/hess-26-1371-2022</ext-link>, 2022. </mixed-citation></ref>
      <ref id="bib1.bibx51"><label>Skaugen and Onof(2014)</label><mixed-citation>Skaugen, T. and Onof, C.: A rainfall-runoff model parameterized from GIS and runoff data, Hydrol. Process., 28, 4529–4542, <ext-link xlink:href="https://doi.org/10.1002/hyp.9968" ext-link-type="DOI">10.1002/hyp.9968</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx52"><label>Tarasova et al.(2019)Tarasova, Merz, Kiss, Basso, Blöschl, Merz, Viglione, Plötner, Guse, Schumann, Fischer, Ahrens, Anwar, Bárdossy, Bühler, Haberlandt, Kreibich, Krug, Lun, Müller-Thomy, Pidoto, Primo, Seidel, Vorogushyn, and Wietzke</label><mixed-citation>Tarasova, L., Merz, R., Kiss, A., Basso, S., Blöschl, G., Merz, B., Viglione, A., Plötner, S., Guse, B., Schumann, A., Fischer, S., Ahrens, B., Anwar, F., Bárdossy, A., Bühler, P., Haberlandt, U., Kreibich, H., Krug, A., Lun, D., Müller-Thomy, H., Pidoto, R., Primo, C., Seidel, J., Vorogushyn, S., and Wietzke, L.: Causative classification of river flood events, WIREs Water, 6, e1353, <ext-link xlink:href="https://doi.org/10.1002/wat2.1353" ext-link-type="DOI">10.1002/wat2.1353</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx53"><label>Tveito(2021)</label><mixed-citation> Tveito, O.: Norwegian standard climate normals 1991–2020 – the methodological approach, Tech. rep., MET report 5/2021, The Norwegian Meteorological Institute, ISSN 2387-4201, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx54"><label>Vormoor et al.(2015)Vormoor, Lawrence, Heistermann, and Bronstert</label><mixed-citation>Vormoor, K., Lawrence, D., Heistermann, M., and Bronstert, A.: Climate change impacts on the seasonality and generation processes of floods – projections and uncertainties for catchments with mixed snowmelt/rainfall regimes, Hydrol. Earth Syst. Sci., 19, 913–931, <ext-link xlink:href="https://doi.org/10.5194/hess-19-913-2015" ext-link-type="DOI">10.5194/hess-19-913-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx55"><label>Vormoor et al.(2016)Vormoor, Lawrence, Schlichting, Wilson, and Wong</label><mixed-citation>Vormoor, K., Lawrence, D., Schlichting, L., Wilson, D., and Wong, W. K.: Evidence for changes in the magnitude and frequency of observed rainfall vs. snowmelt driven floods in Norway, J. Hydrol., 538, 33–48, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2016.03.066" ext-link-type="DOI">10.1016/j.jhydrol.2016.03.066</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx56"><label>Winsvold et al.(2014)Winsvold, Andreassen, and Kienholz</label><mixed-citation>Winsvold, S. H., Andreassen, L. M., and Kienholz, C.: Glacier area and length changes in Norway from repeat inventories, The Cryosphere, 8, 1885–1903, <ext-link xlink:href="https://doi.org/10.5194/tc-8-1885-2014" ext-link-type="DOI">10.5194/tc-8-1885-2014</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx57"><label>Wu et al.(2025)Wu, Zhao, Zhang, Li, Qin, and Li</label><mixed-citation>Wu, X., Zhao, Y., Zhang, W., Li, X., Qin, G., and Li, H.: Probabilistic early warning of flash floods using Monte Carlo simulation and hydrological modelling, Eng. Appl. Comp. Fluid, 19, 2523423, <ext-link xlink:href="https://doi.org/10.1080/19942060.2025.2523423" ext-link-type="DOI">10.1080/19942060.2025.2523423</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx58"><label>Yan et al.(2023)Yan, Zhang, Xiong, Yan, Jiang, Xu, Xiong, Yu, Ma, and Xu</label><mixed-citation>Yan, L., Zhang, L., Xiong, L., Yan, P., Jiang, C., Xu, W., Xiong, B., Yu, K., Ma, Q., and Xu, C.-Y.: Flood Frequency Analysis Using Mixture Distributions in Light of Prior Flood Type Classification in Norway, Remote Sens., 15, <ext-link xlink:href="https://doi.org/10.3390/rs15020401" ext-link-type="DOI">10.3390/rs15020401</ext-link>, 2023.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>The ability of LSTM to model snowmelt versus rainfall generated floods</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Addor and Melsen(2019)</label><mixed-citation>
      
Addor, N. and Melsen, L. A.: Legacy, Rather Than Adequacy, Drives the Selection
of Hydrological Models, Water Resour. Res., 55, 378–390,
<a href="https://doi.org/10.1029/2018WR022958" target="_blank">https://doi.org/10.1029/2018WR022958</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Anderson and Radić(2022)</label><mixed-citation>
      
Anderson, S. and Radić, V.: Evaluation and interpretation of convolutional long short-term memory networks for regional hydrological modelling, Hydrol. Earth Syst. Sci., 26, 795–825, <a href="https://doi.org/10.5194/hess-26-795-2022" target="_blank">https://doi.org/10.5194/hess-26-795-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Andreassen et al.(2022)Andreassen, Nagy, Kjøllmoen, and
Leigh</label><mixed-citation>
      
Andreassen, L. M., Nagy, T., Kjøllmoen, B., and Leigh, J. R.: An inventory
of Norway's glaciers and ice-marginal lakes from 2018–19 Sentinel-2 data,
J. Glaciol., 68, 1085–1106, <a href="https://doi.org/10.1017/jog.2022.20" target="_blank">https://doi.org/10.1017/jog.2022.20</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Bakke et al.(2026)</label><mixed-citation>
      
Bakke, S. J., Barna, D. M., Engeland, K., Kolberg, S. A., and Nordeide, S.: Data for “The ability of LSTM to model snowmelt versus rainfall generated floods”, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.22942372" target="_blank">https://doi.org/10.5281/zenodo.22942372</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Barna et al.(2023a)Barna, Engeland, Kneib,
Thorarinsdottir, and Xu</label><mixed-citation>
      
Barna, D. M., Engeland, K., Kneib, T., Thorarinsdottir, T. L., and Xu, C.-Y.: Regional index flood estimation at multiple durations with generalized additive models, EGUsphere [preprint], <a href="https://doi.org/10.5194/egusphere-2023-2335" target="_blank">https://doi.org/10.5194/egusphere-2023-2335</a>, 2023a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Barna et al.(2023b)Barna, Engeland, Thorarinsdottir, and
Xu</label><mixed-citation>
      
Barna, D. M., Engeland, K., Thorarinsdottir, T. L., and Xu, C.-Y.: Flexible and
consistent Flood–Duration–Frequency modeling: A Bayesian approach,
J. Hydrol., 620, 129448,
<a href="https://doi.org/10.1016/j.jhydrol.2023.129448" target="_blank">https://doi.org/10.1016/j.jhydrol.2023.129448</a>, 2023b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Berghuijs et al.(2019)Berghuijs, Harrigan, Molnar, Slater, and
Kirchner</label><mixed-citation>
      
Berghuijs, W. R., Harrigan, S., Molnar, P., Slater, L. J., and Kirchner, J. W.:
The Relative Importance of Different Flood-Generating Mechanisms Across
Europe, Water Resour. Res., 55, 4582–4593, <a href="https://doi.org/10.1029/2019WR024841" target="_blank">https://doi.org/10.1029/2019WR024841</a>,
2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Bergström(1976)</label><mixed-citation>
      
Bergström, S.: Development and application of a conceptual runoff model for
Scandinavian catchments, Tech. rep., SMHI Report Nr RHO 7/1976, The Swedish
Meteorological and Hydrological Institute (SMHI), ISSN 0347-7827, 1976.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Bocharov(2023)</label><mixed-citation>
      
Bocharov, G.: pyextremes,
<a href="https://github.com/georgebv/pyextremes" target="_blank"/> (last access: 2 December 2025), 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Brunner et al.(2021)Brunner, Melsen, Wood, Rakovec, Mizukami, Knoben,
and Clark</label><mixed-citation>
      
Brunner, M. I., Melsen, L. A., Wood, A. W., Rakovec, O., Mizukami, N., Knoben, W. J. M., and Clark, M. P.: Flood spatial coherence, triggers, and performance in hydrological simulations: large-sample evaluation of four streamflow-calibrated models, Hydrol. Earth Syst. Sci., 25, 105–119, <a href="https://doi.org/10.5194/hess-25-105-2021" target="_blank">https://doi.org/10.5194/hess-25-105-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Doherty(2004)</label><mixed-citation>
      
Doherty, J.: PEST: Model Independent Parameter Estimation, User Manual, 5th edn., <a href="https://www.nrc.gov/docs/ML0923/ML092360221.pdf" target="_blank"/> (last access: 2 October 2026), 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Dyrrdal et al.(2025)Dyrrdal, Bakke, Hanssen-Bauer, Mayer, Nilsen,
Nilsen, Paasche, Saloranta, and Årthun</label><mixed-citation>
      
Dyrrdal, A., Bakke, S., Hanssen-Bauer, I., Mayer, S., Nilsen, I., Nilsen, J.,
Paasche, Ø., Saloranta, T., and Årthun, M. (Eds.): Klima i Norge –
kunnskapsgrunnlag for klimatilpasning oppdatert i 2025, NCCS-rapport 1/2025, The Norwegian Center for Climate Services (NCCS), <a href="https://doi.org/10.60839/4rgq-nn84" target="_blank">https://doi.org/10.60839/4rgq-nn84</a>,
2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Engeland et al.(2016)Engeland, Schlichting, Randen, Nordtun, Reitan,
Wang, Holmqvist, Voksø, and Eide</label><mixed-citation>
      
Engeland, K., Schlichting, L., Randen, F., Nordtun, K., Reitan, T., Wang, T.,
Holmqvist, E., Voksø, A., and Eide, V.: Utvalg og kvalitetssikring av
flomdata forflomfrekvensanalyser, Technical Report 85/2016, The
Norwegian Water Resources and Energy Directorate (NVE), ISBN
978-82-410-1538-0,
<a href="https://publikasjoner.nve.no/rapport/2016/rapport2016_85.pdf" target="_blank"/> (last access: 2 October 2026),
2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Engeland et al.(2020)Engeland, Glad, Hamududu, Li, Reitan, and
Stenius</label><mixed-citation>
      
Engeland, K., Glad, P., Hamududu, B. H., Li, H., Reitan, T., and Stenius,
S. M.: Lokal og regional flomfrekvensanalyse, Technical Report
10/2020, The Norwegian Water Resources and Energy Directorate (NVE), ISBN
978-82-410-2014-8,
<a href="https://publikasjoner.nve.no/rapport/2020/rapport2020_10.pdf" target="_blank"/> (last access: 2 October 2026),
2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Frame et al.(2022)Frame, Kratzert, Klotz, Gauch, Shalev, Gilon,
Qualls, Gupta, and Nearing</label><mixed-citation>
      
Frame, J. M., Kratzert, F., Klotz, D., Gauch, M., Shalev, G., Gilon, O., Qualls, L. M., Gupta, H. V., and Nearing, G. S.: Deep learning rainfall–runoff predictions of extreme events, Hydrol. Earth Syst. Sci., 26, 3377–3392, <a href="https://doi.org/10.5194/hess-26-3377-2022" target="_blank">https://doi.org/10.5194/hess-26-3377-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Gauch et al.(2021)Gauch, Kratzert, Klotz, Nearing, Lin, and
Hochreiter</label><mixed-citation>
      
Gauch, M., Kratzert, F., Klotz, D., Nearing, G., Lin, J., and Hochreiter, S.: Rainfall–runoff prediction at multiple timescales with a single Long Short-Term Memory network, Hydrol. Earth Syst. Sci., 25, 2045–2062, <a href="https://doi.org/10.5194/hess-25-2045-2021" target="_blank">https://doi.org/10.5194/hess-25-2045-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>GeoNorge(2026a)</label><mixed-citation>
      
GeoNorge: Totalnedbørfelt til målestasjon, The Norwegian Mapping Authority (Kartverket),
<a href="https://kartkatalog.geonorge.no/metadata/totalnedboerfelt-til-maalestasjon/ac1c71db-9850-4e89-8162-2baba8b980e7" target="_blank"/>
(last access: 10 February 2026), 2026a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>GeoNorge(2026b)</label><mixed-citation>
      
GeoNorge: ELVIS elvenett, The Norwegian Mapping Authority
(Kartverket),
<a href="https://kartkatalog.geonorge.no/metadata/elvis-elvenett/3f95a194-0968-4457-a500-912958de3d39" target="_blank"/>
(last access: 10 February 2026), 2026b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>GeoNorge(2026c)</label><mixed-citation>
      
GeoNorge: Løsmasser, infiltrasjonEvne, The Norwegian Mapping Authority (Kartverket),
<a href="https://kartkatalog.geonorge.no/metadata/loesmasser/3de4ddf6-d6b8-4398-8222-f5c47791a757" target="_blank"/>, (last access: 10 February 2026), 2026c.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Gers et al.(2000)Gers, Schmidhuber, and Cummins</label><mixed-citation>
      
Gers, F. A., Schmidhuber, J., and Cummins, F.: Learning to Forget: Continual
Prediction with LSTM, Neural Comput., 12, 2451–2471,
<a href="https://doi.org/10.1162/089976600300015015" target="_blank">https://doi.org/10.1162/089976600300015015</a>, 2000.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Gupta et al.(2009)Gupta, Kling, Yilmaz, and
Martinez</label><mixed-citation>
      
Gupta, H. V., Kling, H., Yilmaz, K. K., and Martinez, G. F.: Decomposition of
the mean squared error and NSE performance criteria: Implications for
improving hydrological modelling, J. Hydrol., 377, 80–91,
<a href="https://doi.org/10.1016/j.jhydrol.2009.08.003" target="_blank">https://doi.org/10.1016/j.jhydrol.2009.08.003</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Hagen et al.(2023)Hagen, Hasibi, Leblois, Lawrence, and
Sorteberg</label><mixed-citation>
      
Hagen, J. S., Hasibi, R., Leblois, E., Lawrence, D., and Sorteberg, A.:
Reconstructing daily streamflow and floods from large-scale atmospheric
variables with feed-forward and recurrent neural networks in high latitude
climates, Hydrolog. Sci. J., 68, 412–431,
<a href="https://doi.org/10.1080/02626667.2023.2165927" target="_blank">https://doi.org/10.1080/02626667.2023.2165927</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Hochreiter and Schmidhuber(1997)</label><mixed-citation>
      
Hochreiter, S. and Schmidhuber, J.: Long short-term memory, Neural Compu.,
9, 1735–1780, <a href="https://doi.org/10.1162/neco.1997.9.8.1735" target="_blank">https://doi.org/10.1162/neco.1997.9.8.1735</a>, 1997.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Jiang et al.(2022)Jiang, Bevacqua, and Zscheischler</label><mixed-citation>
      
Jiang, S., Bevacqua, E., and Zscheischler, J.: River flooding mechanisms and their changes in Europe revealed by explainable machine learning, Hydrol. Earth Syst. Sci., 26, 6339–6359, <a href="https://doi.org/10.5194/hess-26-6339-2022" target="_blank">https://doi.org/10.5194/hess-26-6339-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Kartverket(2026)</label><mixed-citation>
      
Kartverket: Høydedata, The Norwegian Mapping Authority (Kartverket), <a href="https://hoydedata.no/LaserInnsyn2/" target="_blank"/>
(last access: 10 February 2026),
2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Killingtveit and Saelthun(1995)</label><mixed-citation>
      
Killingtveit, Å and Saelthun, N. R.: Hydrological models, Hydropower development, Norwegian Inst. of
Technology, Dept. of Hydraulic Engineering, Vol. 7, 99–128, ISBN 978-82-7598-026-5, 1995.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Kingma and Ba(2017)</label><mixed-citation>
      
Kingma, D. P. and Ba, J.: Adam: A Method for Stochastic Optimization. Proceedings of the 3rd International Conference on Learning Representations (ICLR 2015), San Diego, CA, USA, <a href="https://mlanthology.org/iclr/2015/kingma2015iclr-adam/" target="_blank"/> (2 October 2026), 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Kling et al.(2012)Kling, Fuchs, and Paulin</label><mixed-citation>
      
Kling, H., Fuchs, M., and Paulin, M.: Runoff conditions in the upper Danube
basin under an ensemble of climate change scenarios, J. Hydrol.,
424—425, 264–277, <a href="https://doi.org/10.1016/j.jhydrol.2012.01.011" target="_blank">https://doi.org/10.1016/j.jhydrol.2012.01.011</a>, 2012.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Kratzert et al.(2018)Kratzert, Klotz, Brenner, Schulz, and
Herrnegger</label><mixed-citation>
      
Kratzert, F., Klotz, D., Brenner, C., Schulz, K., and Herrnegger, M.: Rainfall–runoff modelling using Long Short-Term Memory (LSTM) networks, Hydrol. Earth Syst. Sci., 22, 6005–6022, <a href="https://doi.org/10.5194/hess-22-6005-2018" target="_blank">https://doi.org/10.5194/hess-22-6005-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Kratzert et al.(2019a)Kratzert, Herrnegger, Klotz,
Hochreiter, and Klambauer</label><mixed-citation>
      
Kratzert, F., Herrnegger, M., Klotz, D., Hochreiter, S., and Klambauer, G.:
NeuralHydrology – Interpreting LSTMs in Hydrology, Springer
International Publishing, Cham, 347–362,, ISBN 978-3-030-28954-6,
<a href="https://doi.org/10.1007/978-3-030-28954-6_19" target="_blank">https://doi.org/10.1007/978-3-030-28954-6_19</a>, 2019a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Kratzert et al.(2019b)Kratzert, Klotz, Shalev,
Klambauer, Hochreiter, and Nearing</label><mixed-citation>
      
Kratzert, F., Klotz, D., Shalev, G., Klambauer, G., Hochreiter, S., and Nearing, G.: Towards learning universal, regional, and local hydrological behaviors via machine learning applied to large-sample datasets, Hydrol. Earth Syst. Sci., 23, 5089–5110, <a href="https://doi.org/10.5194/hess-23-5089-2019" target="_blank">https://doi.org/10.5194/hess-23-5089-2019</a>, 2019b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Kratzert et al.(2021)Kratzert, Klotz, Hochreiter, and
Nearing</label><mixed-citation>
      
Kratzert, F., Klotz, D., Hochreiter, S., and Nearing, G. S.: A note on leveraging synergy in multiple meteorological data sets with deep learning for rainfall–runoff modeling, Hydrol. Earth Syst. Sci., 25, 2685–2703, <a href="https://doi.org/10.5194/hess-25-2685-2021" target="_blank">https://doi.org/10.5194/hess-25-2685-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Kratzert et al.(2022)Kratzert, Gauch, Nearing, and
Klotz</label><mixed-citation>
      
Kratzert, F., Gauch, M., Nearing, G., and Klotz, D.: NeuralHydrology – A
Python library for Deep Learning research in hydrology, Journal of Open
Source Software, 7, 4050, <a href="https://doi.org/10.21105/joss.04050" target="_blank">https://doi.org/10.21105/joss.04050</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Kratzert et al.(2024)Kratzert, Gauch, Klotz, and
Nearing</label><mixed-citation>
      
Kratzert, F., Gauch, M., Klotz, D., and Nearing, G.: HESS Opinions: Never train a Long Short-Term Memory (LSTM) network on a single basin, Hydrol. Earth Syst. Sci., 28, 4187–4201, <a href="https://doi.org/10.5194/hess-28-4187-2024" target="_blank">https://doi.org/10.5194/hess-28-4187-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Lawrence et al.(2009)Lawrence, Haddeland, and
Langsholt</label><mixed-citation>
      
Lawrence, D., Haddeland, I., and Langsholt, E.: Calibration of HBV hydrological
models using PEST parameter estimation, Technical Report 1/2009,
The Norwegian Water Resources and Energy Directorate (NVE), ISBN
78-82-410-0680-7, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Lees et al.(2022)Lees, Reece, Kratzert, Klotz, Gauch, De Bruijn,
Kumar Sahu, Greve, Slater, and Dadson</label><mixed-citation>
      
Lees, T., Reece, S., Kratzert, F., Klotz, D., Gauch, M., De Bruijn, J., Kumar Sahu, R., Greve, P., Slater, L., and Dadson, S. J.: Hydrological concept formation inside long short-term memory (LSTM) networks, Hydrol. Earth Syst. Sci., 26, 3079–3101, <a href="https://doi.org/10.5194/hess-26-3079-2022" target="_blank">https://doi.org/10.5194/hess-26-3079-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Liu et al.(2025)Liu, Shen, O'Donncha, Song, Zhi, Beck, Bindas,
Kraabel, and Lawson</label><mixed-citation>
      
Liu, J., Shen, C., O'Donncha, F., Song, Y., Zhi, W., Beck, H. E., Bindas, T., Kraabel, N., and Lawson, K.: From RNNs to Transformers: benchmarking deep learning architectures for hydrologic prediction, Hydrol. Earth Syst. Sci., 29, 6811–6828, <a href="https://doi.org/10.5194/hess-29-6811-2025" target="_blank">https://doi.org/10.5194/hess-29-6811-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Lussana et al.(2019)Lussana, Tveito, Dobler, and
Tunheim</label><mixed-citation>
      
Lussana, C., Tveito, O. E., Dobler, A., and Tunheim, K.: seNorge_2018, daily precipitation, and temperature datasets over Norway, Earth Syst. Sci. Data, 11, 1531–1551, <a href="https://doi.org/10.5194/essd-11-1531-2019" target="_blank">https://doi.org/10.5194/essd-11-1531-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Martel et al.(2025)Martel, Arsenault, Turcotte, Castañeda Gonzalez,
Brissette, Armstrong, Mailhot, Pelletier-Dumont, Lachance-Cloutier,
Rondeau-Genesse, and Caron</label><mixed-citation>
      
Martel, J.-L., Arsenault, R., Turcotte, R., Castañeda-Gonzalez, M., Brissette, F., Armstrong, W., Mailhot, E., Pelletier-Dumont, J., Lachance-Cloutier, S., Rondeau-Genesse, G., and Caron, L.-P.: Exploring the ability of LSTM-based hydrological models to simulate streamflow time series for flood frequency analysis, Hydrol. Earth Syst. Sci., 29, 4951–4968, <a href="https://doi.org/10.5194/hess-29-4951-2025" target="_blank">https://doi.org/10.5194/hess-29-4951-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>MET Norway(2026)</label><mixed-citation>
      
MET Norway: SeNorge_2018, The Norwegian Meteorological institute (MET
Norway),
<a href="https://thredds.met.no/thredds/catalog/senorge/seNorge_2018/catalog.html" target="_blank"/>
(last access: 11 February 2026), 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Nearing et al.(2024)Nearing, Cohen, Dube, Gauch, Gilon, Harrigan,
Hassidim, Klotz, Kratzert, Metzger et al.</label><mixed-citation>
      
Nearing, G.,
Cohen, D.,
Dube, V.,
Gauch, M.,
Gilon, O.,
Harrigan, S.,
Hassidim, A.,
Klotz, D.,
Kratzert, F.,
Metzger, A.,
Nevo, S.,
Pappenberger, F.,
Prudhomme, C.,
Shalev, G.,
Shenzis, S.,
Tekalign, T. Y.,
Weitzner, D., and
Matias, Y.: Global prediction of
extreme floods in ungauged watersheds, Nature, 627, 559–563,
<a href="https://doi.org/10.1038/s41586-024-07145-1" target="_blank">https://doi.org/10.1038/s41586-024-07145-1</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>NVE(2026a)</label><mixed-citation>
      
NVE: NVE Hydrological API (HydAPI), The Norwegian Water Resources and Energy
Directorate (NVE), <a href="https://hydapi.nve.no/" target="_blank"/>
(last access: 10 February 2026), 2026a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>NVE(2026b)</label><mixed-citation>
      
NVE: Sildre, The Norwegian Water Resources and Energy Directorate (NVE), <a href="https://sildre.nve.no/" target="_blank"/> (last access: 10 February 2026),
2026b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>NVE and MET Norway(2026)</label><mixed-citation>
      
NVE and MET Norway: Snowmelt, The Norwegian Water Resources and Energy
Directorate (NVE) and the Norwegian Meteorological institute (MET Norway), <a href="https://thredds.met.no/thredds/catalog/senorge/seNorge_2018/catalog.html" target="_blank"/> (last access: 10 February 2026),
2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Roksvåg et al.(2026)Roksvåg, Vandeskog, Wulff, and
Wergeland</label><mixed-citation>
      
Roksvåg, T., Vandeskog, S. M., Wulff, C., and Wergeland, K.: An LSTM
network for joint modeling of streamflow and hydropower generation for
run-of-river plants, J. Hydrol., 667, 134890,
<a href="https://doi.org/10.1016/j.jhydrol.2025.134890" target="_blank">https://doi.org/10.1016/j.jhydrol.2025.134890</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Ruan and Langsholt(2017)</label><mixed-citation>
      
Ruan, G. and Langsholt, E.: Rekalibrering av flomvarslingas HBV-modeller med
inndata fra seNorge, versjon 2.0, Technical Report 71/2017, The
Norwegian Water Resources and Energy Directorate (NVE), ISBN
978-82-410-1624-0,
<a href="https://publikasjoner.nve.no/rapport/2017/rapport2017_71.pdf" target="_blank"/> (last access: 2 October 2026),
2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Ruzzante et al.(2026)Ruzzante, Knoben, Wagener, Gleeson, and
Schnorbus</label><mixed-citation>
      
Ruzzante, S. W., Knoben, W. J. M., Wagener, T., Gleeson, T., and Schnorbus, M.: Technical note: High Nash–Sutcliffe Efficiencies conceal poor simulations of interannual variance in seasonal regimes, Hydrol. Earth Syst. Sci., 30, 2337–2355, <a href="https://doi.org/10.5194/hess-30-2337-2026" target="_blank">https://doi.org/10.5194/hess-30-2337-2026</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Sælthun(1996)</label><mixed-citation>
      
Sælthun, N. R.: The Nordic HBV Model, Technical Report 7/1996,
The Norwegian Water Resources and Energy Directorate (NVE), ISBN
82-410-0273-4,
<a href="https://publikasjoner.nve.no/publication/1996/publication1996_07.pdf" target="_blank"/> (last access: 2 October 2026),
1996.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Saloranta(2016)</label><mixed-citation>
      
Saloranta, T. M.: Operational snow mapping with simplified data assimilation
using the seNorge snow model, J. Hydrol., 538, 314–325,
<a href="https://doi.org/10.1016/j.jhydrol.2016.03.061" target="_blank">https://doi.org/10.1016/j.jhydrol.2016.03.061</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Seibert and Bergström(2022)</label><mixed-citation>
      
Seibert, J. and Bergström, S.: A retrospective on hydrological catchment modelling based on half a century with the HBV model, Hydrol. Earth Syst. Sci., 26, 1371–1388, <a href="https://doi.org/10.5194/hess-26-1371-2022" target="_blank">https://doi.org/10.5194/hess-26-1371-2022</a>, 2022.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Skaugen and Onof(2014)</label><mixed-citation>
      
Skaugen, T. and Onof, C.: A rainfall-runoff model parameterized from GIS and
runoff data, Hydrol. Process., 28, 4529–4542, <a href="https://doi.org/10.1002/hyp.9968" target="_blank">https://doi.org/10.1002/hyp.9968</a>,
2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Tarasova et al.(2019)Tarasova, Merz, Kiss, Basso, Blöschl, Merz,
Viglione, Plötner, Guse, Schumann, Fischer, Ahrens, Anwar, Bárdossy,
Bühler, Haberlandt, Kreibich, Krug, Lun, Müller-Thomy, Pidoto, Primo,
Seidel, Vorogushyn, and Wietzke</label><mixed-citation>
      
Tarasova, L., Merz, R., Kiss, A., Basso, S., Blöschl, G., Merz, B.,
Viglione, A., Plötner, S., Guse, B., Schumann, A., Fischer, S., Ahrens,
B., Anwar, F., Bárdossy, A., Bühler, P., Haberlandt, U., Kreibich,
H., Krug, A., Lun, D., Müller-Thomy, H., Pidoto, R., Primo, C., Seidel,
J., Vorogushyn, S., and Wietzke, L.: Causative classification of river flood
events, WIREs Water, 6, e1353, <a href="https://doi.org/10.1002/wat2.1353" target="_blank">https://doi.org/10.1002/wat2.1353</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Tveito(2021)</label><mixed-citation>
      
Tveito, O.: Norwegian standard climate normals 1991–2020 – the methodological
approach, Tech. rep., MET report 5/2021, The Norwegian Meteorological
Institute, ISSN 2387-4201, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Vormoor et al.(2015)Vormoor, Lawrence, Heistermann, and
Bronstert</label><mixed-citation>
      
Vormoor, K., Lawrence, D., Heistermann, M., and Bronstert, A.: Climate change impacts on the seasonality and generation processes of floods – projections and uncertainties for catchments with mixed snowmelt/rainfall regimes, Hydrol. Earth Syst. Sci., 19, 913–931, <a href="https://doi.org/10.5194/hess-19-913-2015" target="_blank">https://doi.org/10.5194/hess-19-913-2015</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Vormoor et al.(2016)Vormoor, Lawrence, Schlichting, Wilson, and
Wong</label><mixed-citation>
      
Vormoor, K., Lawrence, D., Schlichting, L., Wilson, D., and Wong, W. K.:
Evidence for changes in the magnitude and frequency of observed rainfall vs.
snowmelt driven floods in Norway, J. Hydrol., 538, 33–48,
<a href="https://doi.org/10.1016/j.jhydrol.2016.03.066" target="_blank">https://doi.org/10.1016/j.jhydrol.2016.03.066</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Winsvold et al.(2014)Winsvold, Andreassen, and
Kienholz</label><mixed-citation>
      
Winsvold, S. H., Andreassen, L. M., and Kienholz, C.: Glacier area and length changes in Norway from repeat inventories, The Cryosphere, 8, 1885–1903, <a href="https://doi.org/10.5194/tc-8-1885-2014" target="_blank">https://doi.org/10.5194/tc-8-1885-2014</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Wu et al.(2025)Wu, Zhao, Zhang, Li, Qin, and
Li</label><mixed-citation>
      
Wu, X., Zhao, Y., Zhang, W., Li, X., Qin, G., and Li, H.: Probabilistic early
warning of flash floods using Monte Carlo simulation and hydrological
modelling, Eng. Appl. Comp. Fluid, 19,
2523423, <a href="https://doi.org/10.1080/19942060.2025.2523423" target="_blank">https://doi.org/10.1080/19942060.2025.2523423</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Yan et al.(2023)Yan, Zhang, Xiong, Yan, Jiang, Xu, Xiong, Yu, Ma, and
Xu</label><mixed-citation>
      
Yan, L., Zhang, L., Xiong, L., Yan, P., Jiang, C., Xu, W., Xiong, B., Yu, K.,
Ma, Q., and Xu, C.-Y.: Flood Frequency Analysis Using Mixture Distributions
in Light of Prior Flood Type Classification in Norway, Remote Sens., 15,
<a href="https://doi.org/10.3390/rs15020401" target="_blank">https://doi.org/10.3390/rs15020401</a>, 2023.

    </mixed-citation></ref-html>--></article>
