<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">HESS</journal-id><journal-title-group>
    <journal-title>Hydrology and Earth System Sciences</journal-title>
    <abbrev-journal-title abbrev-type="publisher">HESS</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Hydrol. Earth Syst. Sci.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7938</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/hess-30-5343-2026</article-id><title-group><article-title>Process diagnostics of snowmelt runoff in global hydrological and land surface models – Part 1: A systematic evaluation across basins of increasing complexity</article-title><alt-title>Process diagnostics for snowmelt runoff</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Lei</surname><given-names>Xiangyong</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-9971-8615</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Lin</surname><given-names>Haomei</given-names></name>
          
        <ext-link>https://orcid.org/0009-0002-8172-4520</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Zheng</surname><given-names>Kaihao</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-3722-4772</ext-link></contrib>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Lin</surname><given-names>Peirong</given-names></name>
          <email>peironglinlin@pku.edu.cn</email>
        <ext-link>https://orcid.org/0000-0002-7275-7470</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Institute of Remote Sensing and Geographic Information Systems, School of Earth and Space Sciences, Peking University, Beijing, 100871, China</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Peirong Lin (peironglinlin@pku.edu.cn)</corresp></author-notes><pub-date><day>21</day><month>August</month><year>2026</year></pub-date>
      
      <volume>30</volume>
      <issue>16</issue>
      <fpage>5343</fpage><lpage>5372</lpage>
      <history>
        <date date-type="received"><day>5</day><month>December</month><year>2025</year></date>
           <date date-type="rev-request"><day>20</day><month>January</month><year>2026</year></date>
           <date date-type="rev-recd"><day>25</day><month>June</month><year>2026</year></date>
           <date date-type="accepted"><day>1</day><month>August</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Xiangyong Lei et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026.html">This article is available from https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026.html</self-uri><self-uri xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026.pdf">The full text article is available as a PDF file from https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e104">Accurate simulation of snowmelt runoff (SMR) is critical for water resource management. However, despite the abundance of global hydrological models, little is known about their SMR performance. This study presents a comprehensive evaluation of SMR across 15 state-of-the-art large-scale models and runoff products by focusing on their biases in first-order indices, i.e., the total volume (<inline-formula><mml:math id="M1" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), peak flow (<inline-formula><mml:math id="M2" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), and centroid timing (CTQ) of runoff in the snowmelt period. Then by introducing 1455 snow-dominated basins with diverse topography and vegetation complexities, we further proposed a novel model robustness metric to test how different models perform under increasing basin complexity, thereby allowing for a quantification on how they adapt to complex environmental conditions. Our results reveal that (1) most models exhibit underestimated <inline-formula><mml:math id="M3" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M4" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and predict CTQ too early. These biases are particularly pronounced in regions such as the western United States, northern Europe, and northeastern China. (2) Model biases systematically increase with basin complexity, with CTQ exhibiting strong sensitivity to mean elevation and topographic variability, while <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> being shaped more by mean elevation and the diversity of vegetation types in the basin. (3) The robustness assessment further shows that observation-constrained runoff products exhibit the most outstanding performance (i.e., low biases and strong adaptability to stern conditions), followed by the hydrological and land surface models. Notably, while global hydrological models generally exhibit stronger robustness in simulating <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, land surface models show a clear advantage in simulating CTQ, highlighting their structural strength in capturing melt timing rather than runoff magnitude. This study provides a large-sample benchmark for SMR evaluation and complements existing model assessment approaches by examining model performance across basin complexity gradients, offering useful insights for future model development and uncertainty reduction.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>National Key Research and Development Program of China</funding-source>
<award-id>2022YFF0801303</award-id>
</award-group>
<award-group id="gs2">
<funding-source>Beijing Nova Program</funding-source>
<award-id>20230484302</award-id>
</award-group>
<award-group id="gs3">
<funding-source>National Natural Science Foundation of China</funding-source>
<award-id>42371481</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e205">Snowmelt runoff (SMR) is a critical freshwater resource supporting human society and agricultural development. Globally, SMR contributes approximately 50 % of the annual runoff for more than 26 % of the terrestrial land area, directly supplying freshwater for about one-sixth of the world's population <xref ref-type="bibr" rid="bib1.bibx38" id="paren.1"/>. Accurate simulation of SMR is therefore essential for effective water resources management and for assessing the impacts of climate change on water resources. This need becomes increasingly urgent under future climate change scenarios, where established snow–runoff relationships are expected to shift substantially <xref ref-type="bibr" rid="bib1.bibx47" id="paren.2"/>, necessitating a systematic understanding of SMR dynamics.</p>
      <p id="d2e214">However, although large scale hydrological models have been widely used to simulate SMR in cold region hydrology, accurately reproducing key SMR characteristics across diverse basin conditions remains challenging. This challenge is partly related to the coupled nature of SMR-related processes, including rainfall–snowfall partitioning, sublimation, canopy interception, and snowmelt dynamics, as well as to uncertainties in forcing data, parameterization <xref ref-type="bibr" rid="bib1.bibx24" id="paren.3"/>, and model structure. Such a challenge is consistently highlighted by multi-Model Intercomparison Projects (MIPs). For example, <xref ref-type="bibr" rid="bib1.bibx21" id="text.4"/> reported substantially lower model skill in cold regions compared to non-cold regions. Similarly, <xref ref-type="bibr" rid="bib1.bibx14" id="text.5"/> revealed markedly larger runoff biases in cold regions, particularly for extreme events. However, existing model assessments have largely emphasized the overall model performance at annual or seasonal time scales <xref ref-type="bibr" rid="bib1.bibx44" id="paren.6"/>, and comprehensive evaluations explicitly targeting SMR remain rare. These studies generally attributed errors to simplified process representations, parameter uncertainties, or model structures <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx16" id="paren.7"/>, without a direct diagnosis of the underlying processes. As a result, the critical deficiencies in SMR processes are usually entangled with errors from other hydrological processes <xref ref-type="bibr" rid="bib1.bibx6" id="paren.8"/>, compromising our understanding of the behavior and limitations of existing models. This highlights an urgent need for a tailored diagnosis for the SMR processes to guide future model development.</p>
      <p id="d2e236">One of the key factors of a systematic assessment is to identify the appropriate evaluation perspectives. Previous model evaluations have primarily focused on aggregated metrics such as the Nash–Sutcliffe Efficiency (NSE; <xref ref-type="bibr" rid="bib1.bibx33" id="altparen.9"/>) and the Kling–Gupta Efficiency (KGE; <xref ref-type="bibr" rid="bib1.bibx15" id="altparen.10"/>). While they are useful for assessing the overall temporal dynamics, their aggregated nature makes it difficult to disentangle the specific deficiencies tied to process representations. In fact, a satisfying model should accurately capture the major characteristics of the SMR hydrograph, namely its total volume (<inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), peak flow (<inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), and centroid timing (CTQ), which are the first-order metrics crucial to water resource management and utilization. However, past assessments have not been structured to explicitly quantify these dimensions, which may hide the potential trade-offs of certain models (e.g., a model excelling in timing but failing in volume).</p>
      <p id="d2e267">Beyond the major characteristics of a hydrograph, a more critical perspective is to assess how well a model performs across diverse land conditions. This is particularly relevant for SMR simulation, where land surface complexity is known to challenge model accuracy. The complexity is generally related to terrain and vegetation conditions <xref ref-type="bibr" rid="bib1.bibx26 bib1.bibx12" id="paren.11"/>: high-elevation environments introduce intricate rainfall–snowfall partitioning thresholds and enhanced sublimation <xref ref-type="bibr" rid="bib1.bibx40 bib1.bibx42" id="paren.12"/>, while varied topography and vegetation demand sophisticated representations of canopy interception, radiation transfer, and runoff generation. Taken together, a model that performs well or maintains its skill even in the face of increasingly complex land surface conditions should be considered robust. However, past assessments have rarely been designed to test this robustness and reveal which type of models excel under varying levels of complexity. This presents another gap in benchmarking our current modeling capabilities.</p>
      <p id="d2e277">To address these gaps, we gathered 15 state-of-the-art large-scale runoff models and data products across 1455 snow-dominated basins worldwide for the period 1979–2019. Based on this dataset, we systematically evaluated their SMR performance by explicitly considering the three primary characteristics of the SMR hydrograph, followed by a further test of model robustness using a novel metrics to describe the performance under varying land complexities. The models and products include six Inter-Sectoral Impact Model Intercomparison Project (ISIMIP2a) water-sector models (PCR-GLOBWB, DBH, VIC, MATSIRO, CLM40, LPJML), seven ISIMIP3a water-sector models (CWATM, H08, HYDROPY, JULES-W2, MIROC-INTEG-LAND, ORCHIDEE-MICT, WATERGAP2-2E), and two recent global river-flow data products, namely Global Reach-Level A Priori Discharge Estimates for Surface Water and Ocean Topography (GRADES; <xref ref-type="bibr" rid="bib1.bibx29" id="altparen.13"/>) and Global River Discharge Reanalysis (GRDR; <xref ref-type="bibr" rid="bib1.bibx10" id="altparen.14"/>). Our model selection criteria are twofold: first, it should cover a sufficiently wide spectrum of models to facilitate a discussion on the performance and process disparities of different modelling schemes <xref ref-type="bibr" rid="bib1.bibx7 bib1.bibx14 bib1.bibx21" id="paren.15"/>; second, it should cover observation-constrained runoff products, allowing for an assessment of the potential performance gains from gauge or satellite data. Our analysis begins with a systematic assessment of the primary characteristics of SMR (total volume, peak flow, and centroid timing), followed by an analysis of performance stratified by basin complexity and a detailed discussion of model robustness. By combining multiple SMR characteristics with model robustness analysis across basin-complexity gradients, this study aims to provide a complementary perspective for diagnosing model deficiencies and informing future model development and uncertainty reduction.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Data and Methods</title>
<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Data</title>
<sec id="Ch1.S2.SS1.SSS1">
  <label>2.1.1</label><title>ISIMIP2a/3a water sector models and global runoff products</title>
      <p id="d2e311">Table <xref ref-type="table" rid="T1"/> summarizes the 13 state-of-the-art macro-scale water sector models and two data products utilized in this study. Among these, six models belong to ISIMIP2a, while seven are part of ISIMIP3a. All models were obtained from the ISIMIP water sector data repository (<uri>https://data.isimip.org/</uri>, last access: 19 November 2024). For each model, their total runoff (i.e., sum of surface and subsurface runoff) was extracted for river routing (Sect. 2.2), and the key SMR indices were subsequently calculated for evaluation (Sect. 2.3).</p>
      <p id="d2e319">These models were categorized into three groups (see Table <xref ref-type="table" rid="T1"/> and references therein): i.e., six global hydrological models (GHMs: PCR-GLOBWB, VIC, CWATM, H08, HYDROPY, and WATERGAP2-2E), six land surface models (LSMs: DBH, MATSIRO, CLM40, JULES-W2, MIROC-INTEG-LAND, and ORCHIDEE-MICT), and one dynamic global vegetation model (DGVM: LPJML). We categorized the models into these groups primarily because GHMs tend to focus on water balance representation, LSMs are generally more advanced in simulating energy-exchange processes, and DGVMs are better suited for capturing vegetation ecosystem dynamics. Thus, comparative analyses across different model categories may provide insights into their relative strengths and limitations. We also included two observation-constrained datasets for evaluation – GRADES, a global runoff product based on VIC model simulations followed by bias correction using gauge-extrapolated information, and GRDR, a recently released global runoff product that enhances accuracy by assimilating river-width data from Landsat into a model-based discharge simulation framework. Incorporating these two discharge products allows for discussions on the gains brought by observational constraints.</p>
      <p id="d2e324">Because parameter calibration can influence runoff simulation, especially for runoff magnitude and peak flow, we further documented the calibration status of the evaluated models based on the available model documentation and calibration information. Five models, including PCR-GLOBWB, LPJML, HYDROPY, CWATM, and WATERGAP2-2E, were classified as calibrated models, whereas MATSIRO, DBH, CLM40, MIROC-INTEG-LAND, VIC, ORCHIDEE-MICT, JULES-W2, and H08 were grouped as uncalibrated models because no explicit calibration entry was available in the original model information.</p>
      <p id="d2e327">All models were run at a spatial resolution of <inline-formula><mml:math id="M11" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.5</mml:mn><mml:mi mathvariant="italic">°</mml:mi></mml:mrow></mml:math></inline-formula> with daily time steps. Models with similar meteorological forcing and simulation scenarios were purposefully chosen such that our comparisons more directly focus on process diagnostics, although other uncertainty sources such as calibration differences, forcing uncertainty, and model-specific implementation choices cannot be fully eliminated. All models and datasets were matched to the same geospatial framework (i.e., the MERIT-Basins river network; <xref ref-type="bibr" rid="bib1.bibx29" id="altparen.16"/>) to ensure consistency. Further details are provided in Sect. 2.2.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e347"><italic>Overview of models and data products considered in this study.</italic> Calibration status indicates whether explicit calibration information is available from the original model documentation or related model information. Models without explicit calibration information are grouped as uncalibrated in this study.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="6">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="left"/>
     <oasis:colspec colnum="6" colname="col6" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Experiment</oasis:entry>
         <oasis:entry colname="col2">Model</oasis:entry>
         <oasis:entry colname="col3">Climate forcing</oasis:entry>
         <oasis:entry colname="col4">Model class</oasis:entry>
         <oasis:entry colname="col5">Calibration status</oasis:entry>
         <oasis:entry colname="col6">Reference</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">ISIMIP2a</oasis:entry>
         <oasis:entry colname="col2">PCR-GLOBWB</oasis:entry>
         <oasis:entry colname="col3">GSWP3</oasis:entry>
         <oasis:entry colname="col4">GHM</oasis:entry>
         <oasis:entry colname="col5">Calibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx43" id="text.17"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">DBH</oasis:entry>
         <oasis:entry colname="col3">GSWP3</oasis:entry>
         <oasis:entry colname="col4">LSM</oasis:entry>
         <oasis:entry colname="col5">Uncalibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx45" id="text.18"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">VIC</oasis:entry>
         <oasis:entry colname="col3">GSWP3</oasis:entry>
         <oasis:entry colname="col4">GHM</oasis:entry>
         <oasis:entry colname="col5">Uncalibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx27" id="text.19"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">MATSIRO</oasis:entry>
         <oasis:entry colname="col3">GSWP3</oasis:entry>
         <oasis:entry colname="col4">LSM</oasis:entry>
         <oasis:entry colname="col5">Uncalibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx36" id="text.20"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">CLM40</oasis:entry>
         <oasis:entry colname="col3">GSWP3</oasis:entry>
         <oasis:entry colname="col4">LSM</oasis:entry>
         <oasis:entry colname="col5">Uncalibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx34" id="text.21"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">LPJML</oasis:entry>
         <oasis:entry colname="col3">GSWP3</oasis:entry>
         <oasis:entry colname="col4">DGVM</oasis:entry>
         <oasis:entry colname="col5">Calibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx39" id="text.22"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">ISIMIP3a</oasis:entry>
         <oasis:entry colname="col2">CWATM</oasis:entry>
         <oasis:entry colname="col3">GSWP3-W5E5</oasis:entry>
         <oasis:entry colname="col4">GHM</oasis:entry>
         <oasis:entry colname="col5">Calibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx5" id="text.23"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">H08</oasis:entry>
         <oasis:entry colname="col3">GSWP3-W5E5</oasis:entry>
         <oasis:entry colname="col4">GHM</oasis:entry>
         <oasis:entry colname="col5">Uncalibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx18" id="text.24"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">HYDROPY</oasis:entry>
         <oasis:entry colname="col3">GSWP3-W5E5</oasis:entry>
         <oasis:entry colname="col4">GHM</oasis:entry>
         <oasis:entry colname="col5">Calibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx41" id="text.25"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">JULES-W2</oasis:entry>
         <oasis:entry colname="col3">GSWP3-W5E5</oasis:entry>
         <oasis:entry colname="col4">LSM</oasis:entry>
         <oasis:entry colname="col5">Uncalibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx4" id="text.26"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">MIROC-INTEG-LAND</oasis:entry>
         <oasis:entry colname="col3">GSWP3-W5E5</oasis:entry>
         <oasis:entry colname="col4">LSM</oasis:entry>
         <oasis:entry colname="col5">Uncalibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx51" id="text.27"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">ORCHIDEE-MICT</oasis:entry>
         <oasis:entry colname="col3">GSWP3-W5E5</oasis:entry>
         <oasis:entry colname="col4">LSM</oasis:entry>
         <oasis:entry colname="col5">Uncalibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx13" id="text.28"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">WATERGAP2-2E</oasis:entry>
         <oasis:entry colname="col3">GSWP3-W5E5</oasis:entry>
         <oasis:entry colname="col4">GHM</oasis:entry>
         <oasis:entry colname="col5">Calibrated</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx32" id="text.29"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Dataset</oasis:entry>
         <oasis:entry colname="col2">GRADES</oasis:entry>
         <oasis:entry colname="col3">MSWEP</oasis:entry>
         <oasis:entry colname="col4">/</oasis:entry>
         <oasis:entry colname="col5">/</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx29" id="text.30"/>
                    </oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2">GRDR</oasis:entry>
         <oasis:entry colname="col3">MSWEP and ECMWF</oasis:entry>
         <oasis:entry colname="col4">/</oasis:entry>
         <oasis:entry colname="col5">/</oasis:entry>
         <oasis:entry colname="col6">
                      <xref ref-type="bibr" rid="bib1.bibx10" id="text.31"/>
                    </oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S2.SS1.SSS2">
  <label>2.1.2</label><title>Routing runoff with the RAPID vector-based routing model</title>
      <p id="d2e771">To ensure that all models are comparable under the same geospatial framework, we first routed the gridded runoff of all ISIMIP2a and ISIMIP3a models through the same river routing model, RAPID (the Routing Application for Parallel computatIon of Discharge; <xref ref-type="bibr" rid="bib1.bibx8" id="altparen.32"/>). RAPID is an efficient vector-based river routing model that enables intercomparison of discharge at global scales <xref ref-type="bibr" rid="bib1.bibx8" id="paren.33"/>, making it an ideal choice for routing ISIMIP runoff.</p>
      <p id="d2e780">The river network used for routing is MERIT-Basins <xref ref-type="bibr" rid="bib1.bibx29" id="paren.34"/>, a high-resolution, vector-based hydrography dataset constructed from the MERIT-Hydro DEM <xref ref-type="bibr" rid="bib1.bibx48" id="paren.35"/>. We use the area-weighted mapping technique <xref ref-type="bibr" rid="bib1.bibx28" id="paren.36"/> to map the gridded runoff (<inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.5</mml:mn><mml:mi mathvariant="italic">°</mml:mi></mml:mrow></mml:math></inline-formula>) onto the vectorized hydrography to determine lateral inflows, with the connectivity of the river-network topology pre-specified as input. RAPID employs the Muskingum method, which requires two parameters: a weighting factor <inline-formula><mml:math id="M13" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> and the flood-wave travel time <inline-formula><mml:math id="M14" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>. Since <inline-formula><mml:math id="M15" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula> is less sensitive in the Muskingum method, it is typically set to 0.3 globally. In contrast, <inline-formula><mml:math id="M16" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> plays a key role in routing performance and is estimated using river-specific characteristics. Specifically, <inline-formula><mml:math id="M17" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> is calculated for each river reach by combining channel length with flow celerity estimated from Manning's equation, as expressed in Eq. (<xref ref-type="disp-formula" rid="Ch1.E1"/>):

              <disp-formula id="Ch1.E1" content-type="numbered"><label>1</label><mml:math id="M18" display="block"><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mn mathvariant="normal">3</mml:mn><mml:mi>ln⁡</mml:mi></mml:mrow><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:msqrt><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:msqrt><mml:msup><mml:mfenced close=")" open="("><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mi>w</mml:mi><mml:mi>d</mml:mi></mml:mrow><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mi>d</mml:mi><mml:mo>+</mml:mo><mml:mi>w</mml:mi></mml:mrow></mml:mfrac></mml:mstyle></mml:mfenced><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>/</mml:mo><mml:mn mathvariant="normal">3</mml:mn></mml:mrow></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

            Where <inline-formula><mml:math id="M19" display="inline"><mml:mi>l</mml:mi></mml:math></inline-formula> is the length of the river channel, <inline-formula><mml:math id="M20" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is Manning's roughness coefficient and is typically set to <inline-formula><mml:math id="M21" display="inline"><mml:mn mathvariant="normal">0.035</mml:mn></mml:math></inline-formula> for natural rivers <xref ref-type="bibr" rid="bib1.bibx29" id="paren.37"/>. <inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> is the channel slope, and <inline-formula><mml:math id="M23" display="inline"><mml:mi>w</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M24" display="inline"><mml:mi>d</mml:mi></mml:math></inline-formula> are the river width and depth estimated by multi-year average runoff using a long established power-law equation <xref ref-type="bibr" rid="bib1.bibx1" id="paren.38"/>. Note that the choice of river routing model and its parameters can influence both the timing and magnitude of SMR, particularly for <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and CTQ, which are more sensitive to runoff concentration pathways and travel-time estimates than <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. In addition, a scale mismatch exists between the <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mn mathvariant="normal">0.5</mml:mn><mml:mi mathvariant="italic">°</mml:mi></mml:mrow></mml:math></inline-formula> gridded runoff outputs and the finer MERIT Basins river network used for routing. Although the area weighted mapping method and the consistent RAPID routing framework improve comparability across models, they cannot fully resolve subgrid heterogeneity or the allocation of coarse grid runoff to fine river reaches. This mismatch may introduce uncertainty into the absolute evaluation of basin scale runoff characteristics. Nevertheless, because our analysis focuses primarily on long-term mean SMR characteristics and applies the same routing procedure to all models and datasets, we expect this uncertainty to have a limited influence on the relative inter-model comparison. In summary, we standardized the forcing data, geospatial framework, and routing scheme as much as possible to improve inter-model comparability and to focus the evaluation more directly on differences in land model parameterizations and process representations. Nevertheless, forcing and routing-related uncertainties may still affect the simulated SMR characteristics, particularly <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and CTQ, and thus should be considered when interpreting both the absolute model performance and inter-model differences.</p>
</sec>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Methods</title>
<sec id="Ch1.S2.SS2.SSS1">
  <label>2.2.1</label><title>Definition of key SMR characteristics</title>
      <p id="d2e1006">As briefly discussed in Sect. 1, diagnosing processes related to SMR can be challenging due to the many processes involved (Fig. <xref ref-type="fig" rid="F1"/>a). To simplify this, we focus on three first-order characteristics–total runoff (<inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), peak flow (<inline-formula><mml:math id="M30" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), and the centroid timing (CTQ) of runoff in the snowmelt period. <inline-formula><mml:math id="M31" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is closely linked to water availability, <inline-formula><mml:math id="M32" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> determines flood hazard potential, and CTQ provides essential information for water resource management. More importantly, these metrics are directly linked with snow accumulation and melt processes (Fig. <xref ref-type="fig" rid="F1"/>b). Specifically, rainfall–snowfall partitioning controls the precipitation phase and the magnitude of snow accumulation, while snow interception and sublimation regulate snow redistribution and loss. Together, these processes determine the amount of meltable snow, thereby influencing both <inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M34" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. In comparison, the melt process, being more closely linked to energy dynamics, governs the timing and rate of SMR and exerts a stronger control on CTQ (Fig. <xref ref-type="fig" rid="F1"/>b). Compared with a full time-series analysis, our focus on these key runoff characteristics offers a more direct and physically interpretable evaluation of model performance.</p>
      <p id="d2e1082">To obtain these metrics, we first defined the snowmelt period using snow water equivalent (SWE) data from the fifth-generation atmospheric reanalysis of the European Centre for Medium-Range Weather Forecasts (ERA5; <xref ref-type="bibr" rid="bib1.bibx20" id="altparen.39"/>). Specifically, the snowmelt period was defined as the interval from the date of maximum SWE to the date when SWE dropped below 1 mm (Fig. <xref ref-type="fig" rid="F1"/>c). This definition was designed to identify the dominant seasonal snowmelt period in a consistent manner across all basins, rather than to capture every short-term snowmelt-related runoff event. Therefore, it may not fully represent complex melt dynamics such as multi-peak melt seasons, intermittent melt–refreeze cycles, or rain-on-snow events. Here, ERA5 SWE was used mainly to identify the timing window of the main snowmelt period rather than to evaluate SWE magnitude. We acknowledge that ERA5 SWE contains uncertainties, especially in complex terrain where snow accumulation and melt are affected by elevation gradients, slope, aspect, vegetation cover, and sub-grid snow redistribution processes. However, their influence on the extracted SMR characteristics is expected to be limited, because this study focuses on the dominant seasonal snowmelt signal rather than short-term melt events or daily SWE variability. Moreover, <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M36" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ were derived as long-term mean characteristics across multiple years, which reduces the sensitivity of the results to occasional errors in snowmelt-period boundaries. Therefore, although ERA5 SWE uncertainty should be considered, it is unlikely to substantially alter the main runoff characteristics used in this global scale evaluation.</p>
      <p id="d2e1112">After this, <inline-formula><mml:math id="M37" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is defined as the total runoff in the snowmelt periods. <inline-formula><mml:math id="M38" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the maximum discharge in the snowmelt periods. CTQ refers to the calendar date on which cumulative runoff reaches 50 % of the total runoff in the snowmelt period, which measures the timing of concentrated runoff during snowmelt, reflecting both the onset and the rate of melt. These definitions are conceptually consistent with previous studies that examined annual <xref ref-type="bibr" rid="bib1.bibx17" id="paren.40"/> or cold-season runoff <xref ref-type="bibr" rid="bib1.bibx9" id="paren.41"/>, but they are further restricted to the snowmelt period to increase the SMR signals. We also introduced a few filtering steps to ensure the SMR signals are the dominant ones, which will be introduced in Sect. 2.2.2.</p>

      <fig id="F1" specific-use="star"><label>Figure 1</label><caption><p id="d2e1146"><italic>Major physical processes influencing snowmelt runoff (SMR), key runoff characteristics, and the study area.</italic>
<bold>(a)</bold> dominant physical processes during snow accumulation and melt; <bold>(b)</bold> linkages between physical processes and key runoff characteristics; <bold>(c)</bold> definition of the snowmelt period and the three key characteristics used for model evaluation; <bold>(d)</bold> spatial distribution of hydrological stations and the ratio of snowmelt runoff.</p></caption>
            <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f01.png"/>

          </fig>

</sec>
<sec id="Ch1.S2.SS2.SSS2">
  <label>2.2.2</label><title>Catchment selection</title>
      <p id="d2e1177">We obtained daily streamflow data from 1455 gauges (Fig. <xref ref-type="fig" rid="F1"/>d) for evaluation. The gauge selection and associated catchment information were based on the Global Streamflow Characteristics, Hydrometeorology, and Catchment Attributes dataset (GSHA; <xref ref-type="bibr" rid="bib1.bibx50" id="altparen.42"/>), which covers 21 568 gauges. We note that the public GSHA release provides annual and monthly streamflow indices, while the daily discharge time series were obtained through following the GSHA data retrieving packages, and the data were used to calculate <inline-formula><mml:math id="M39" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ. The gauges were then filtered to include only those above <inline-formula><mml:math id="M41" display="inline"><mml:mn mathvariant="normal">30</mml:mn></mml:math></inline-formula>° N (i.e., mid- to high-latitude regions) and those with long-term snow water equivalent (SWE) <inline-formula><mml:math id="M42" display="inline"><mml:mrow><mml:mo>≥</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> mm in at least one month of the cold season (i.e., during October to March). Additionally, to minimize the impact of other processes such as human activities, glacier melt, and permafrost thaw on runoff, we excluded basins with urban land cover <inline-formula><mml:math id="M43" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> %, degree of regulation (DOR) <inline-formula><mml:math id="M44" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> %, a combined fraction of urban and cropland areas <inline-formula><mml:math id="M45" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> %, and glacier or permafrost coverage <inline-formula><mml:math id="M46" display="inline"><mml:mrow><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> %. To identify snowmelt dominated basins, we estimated the relative importance of SMR during the snowmelt period using a residual water-balance proxy (Eq. <xref ref-type="disp-formula" rid="Ch1.E2"/>). Only basins with <inline-formula><mml:math id="M47" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SMR</mml:mi><mml:mi mathvariant="normal">ratio</mml:mi></mml:msub><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula> were retained, indicating that snowmelt signals are likely to dominate runoff during the defined snowmelt period.

              <disp-formula id="Ch1.E2" content-type="numbered"><label>2</label><mml:math id="M48" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SMR</mml:mi><mml:mi mathvariant="normal">ratio</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mi mathvariant="normal">SMR</mml:mi><mml:mi>Q</mml:mi></mml:mfrac></mml:mstyle></mml:mrow></mml:math></disp-formula>

            where SMR was approximated using a residual water-balance approach, following Eq. (<xref ref-type="disp-formula" rid="Ch1.E3"/>):

              <disp-formula id="Ch1.E3" content-type="numbered"><label>3</label><mml:math id="M49" display="block"><mml:mrow><mml:mi mathvariant="normal">SMR</mml:mi><mml:mo>=</mml:mo><mml:mi>Q</mml:mi><mml:mo>-</mml:mo><mml:mi>P</mml:mi><mml:mo>+</mml:mo><mml:mi mathvariant="normal">ET</mml:mi></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M50" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula> was derived from gauge observations in GSHA <xref ref-type="bibr" rid="bib1.bibx50" id="paren.43"/>, <inline-formula><mml:math id="M51" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> from Multi-Source Weighted-Ensemble Precipitation <xref ref-type="bibr" rid="bib1.bibx3" id="paren.44"/>, and ET from the Global Land Evaporation Amsterdam Model <xref ref-type="bibr" rid="bib1.bibx31" id="paren.45"/>.</p>
      <p id="d2e1348">We note that this residual estimate is used only as a pragmatic screening proxy for identifying snowmelt dominated basins, rather than as a rigorous quantification of the exact snowmelt runoff contribution. This approach does not explicitly account for changes in catchment storage, delayed groundwater release, or rainfall–snowmelt interactions, all of which may affect the estimated SMR contribution. Therefore, <inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SMR</mml:mi><mml:mi mathvariant="normal">ratio</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> should be interpreted as an indicator of dominant snowmelt influence during the snowmelt period.</p>
      <p id="d2e1362">Above all, only basins with more than 10 years of runoff observations during 1979–2019 were included. For each year, the proportion of missing daily records had to be <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:mo>&lt;</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula> %, and missing values were filled using linear interpolation. Via the above constraints, potential interferences on basin-scale SMR were minimized, ensuring the robustness of the SMR assessment.</p>
      <p id="d2e1375">As illustrated in Fig. <xref ref-type="fig" rid="F2"/>, the key characteristics of snowmelt runoff exhibit pronounced spatial variations across the Northern Hemisphere. In the western coastal and mountainous regions of the United States, the northeastern United States, northern Europe, and northern Japan, <inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SWE</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> values are relatively high (Fig. <xref ref-type="fig" rid="F2"/>a). This corresponds to later melt onset (<inline-formula><mml:math id="M55" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Melt</mml:mi><mml:mi mathvariant="normal">start</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and melt completion (<inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Melt</mml:mi><mml:mi mathvariant="normal">end</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) in the season, as well as higher <inline-formula><mml:math id="M57" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M58" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> values. The CTQ date is correspondingly later (Fig. <xref ref-type="fig" rid="F2"/>a–f). In comparison, the central United States, the eastern coastal United States, western Europe, and northeastern China generally exhibit lower values of these snowmelt characteristics, indicating relatively less snow, earlier melt timing, and reduced meltwater contributions.</p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e1443"><italic>The spatial pattern of snow and runoff characteristics.</italic>
<bold>(a)</bold> shows maximum snow water equivalent, <bold>(b)</bold> and <bold>(c)</bold> show the start and end timing of snowmelt. <bold>(d)</bold>–<bold>(e)</bold> present total runoff (<inline-formula><mml:math id="M59" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), peak flow (<inline-formula><mml:math id="M60" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>) and centroid timing of runoff (CTQ) during the snowmelt period, respectively. <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Melt</mml:mi><mml:mi mathvariant="normal">start</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M62" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Melt</mml:mi><mml:mi mathvariant="normal">end</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ are shown as the months instead of 'Day of Water Year' (DOY) for better interpretation of the results.</p></caption>
            <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f02.png"/>

          </fig>

</sec>
<sec id="Ch1.S2.SS2.SSS3">
  <label>2.2.3</label><title>Quantifying the complexity of basins</title>
      <p id="d2e1522">To quantify the impact of land surface complexity on SMR simulation, it is essential to first identify the primary factors that may challenge model performance. Based on existing literature <xref ref-type="bibr" rid="bib1.bibx26 bib1.bibx37 bib1.bibx46 bib1.bibx19" id="paren.46"/>, we focus on two key dimensions of basin complexity: <italic>topography</italic> and <italic>vegetation</italic>. Topographic complexity affects SMR simulation through several mechanisms. Mean elevation influences temperature, precipitation phase, snow accumulation, sublimation, and snow–radiation interactions, whereas topographic relief introduces finer-scale variations in surface energy balance, snow redistribution, melt timing, and runoff generation. Vegetation complexity also affects SMR-related processes through canopy density and vegetation composition. Higher canopy density can enhance snow interception, sublimation, and canopy shading, while greater diversity in plant functional types increases landscape heterogeneity in energy exchange and hydrological responses <xref ref-type="bibr" rid="bib1.bibx26" id="paren.47"/>.</p>
      <p id="d2e1537">Correspondingly, we selected four metrics to represent these two major dimensions of basin complexity: mean elevation (denoted as <inline-formula><mml:math id="M63" display="inline"><mml:mi mathvariant="normal">DEM</mml:mi></mml:math></inline-formula>), topographic variability (denoted as <inline-formula><mml:math id="M64" display="inline"><mml:mi mathvariant="normal">DEMstd</mml:mi></mml:math></inline-formula>), mean leaf area index (denoted as <inline-formula><mml:math id="M65" display="inline"><mml:mi mathvariant="normal">LAI</mml:mi></mml:math></inline-formula>), and the entropy of plant functional type (denoted as <inline-formula><mml:math id="M66" display="inline"><mml:mi mathvariant="normal">PFTh</mml:mi></mml:math></inline-formula>). Specifically, <inline-formula><mml:math id="M67" display="inline"><mml:mi mathvariant="normal">DEM</mml:mi></mml:math></inline-formula> represents the mean elevation background and associated climatic controls, while <inline-formula><mml:math id="M68" display="inline"><mml:mi mathvariant="normal">DEMstd</mml:mi></mml:math></inline-formula> reflects within-basin terrain heterogeneity. <inline-formula><mml:math id="M69" display="inline"><mml:mi mathvariant="normal">LAI</mml:mi></mml:math></inline-formula> represents vegetation density and its effects on canopy–snow interactions, whereas <inline-formula><mml:math id="M70" display="inline"><mml:mi mathvariant="normal">PFTh</mml:mi></mml:math></inline-formula> describes vegetation-type diversity and the heterogeneity of vegetation-related hydrological and energy-exchange processes. Together, these metrics form a four-dimensional vector describing the overall topographic and vegetation complexity of a basin, where higher values generally denote more challenging conditions for SMR simulation.</p>
      <p id="d2e1597">To synthesize this multi-faceted complexity, we introduced a basin Complexity Index (CI) by summing the normalized values of these four metrics (Eq. <xref ref-type="disp-formula" rid="Ch1.E4"/>). This index is designed to provide a simple and physically interpretable representation of the major topographic and vegetation controls on SMR simulation difficulty:

              <disp-formula id="Ch1.E4" content-type="numbered"><label>4</label><mml:math id="M71" display="block"><mml:mrow><mml:mi mathvariant="normal">CI</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mi mathvariant="normal">DEM</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">DEMstd</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">LAI</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mi mathvariant="normal">PFTh</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e1639">We acknowledge that these four components may not contribute equally to hydrological complexity and that dependencies may exist among them. Therefore, we further examined the correlations among the CI components (Appendix Fig. <xref ref-type="fig" rid="FA1"/>). This additional analysis shows that, although some dependence exists among the components, especially between <inline-formula><mml:math id="M72" display="inline"><mml:mi mathvariant="normal">DEM</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M73" display="inline"><mml:mi mathvariant="normal">DEMstd</mml:mi></mml:math></inline-formula>, the four variables capture complementary topographic and vegetation controls on SMR simulation.</p>
      <p id="d2e1659">To further clarify the physical meaning of CI, we examined the spatial distributions of both the individual components and the integrated CI (Fig. <xref ref-type="fig" rid="F3"/>). The four components show distinct but complementary spatial patterns, indicating that basin complexity is jointly shaped by elevation background, topographic variability, vegetation density, and vegetation-type heterogeneity. The integrated CI highlights several regions with relatively high basin complexity, including the western United States, western Europe, Japan, and northeastern China, where complex terrain and/or vegetation conditions may pose greater challenges for SMR simulation. These spatial patterns support the use of CI as a synthetic descriptor of basin complexity and provide a basis for interpreting subsequent model-performance changes along the CI gradient.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e1666"><italic>Spatial distribution of basin complexity components and representative examples of the integrated Complexity Index.</italic> Panels show <bold>(a)</bold> <inline-formula><mml:math id="M74" display="inline"><mml:mi mathvariant="normal">DEM</mml:mi></mml:math></inline-formula>, <bold>(b)</bold> <inline-formula><mml:math id="M75" display="inline"><mml:mi mathvariant="normal">DEMstd</mml:mi></mml:math></inline-formula>, <bold>(c)</bold> <inline-formula><mml:math id="M76" display="inline"><mml:mi mathvariant="normal">LAI</mml:mi></mml:math></inline-formula>, <bold>(d)</bold> <inline-formula><mml:math id="M77" display="inline"><mml:mi mathvariant="normal">PFTh</mml:mi></mml:math></inline-formula>, and <bold>(e)</bold> the integrated Complexity Index (CI) across the study basins. The three starred basins in panel <bold>(e)</bold> are representative examples of Low CI, Medium CI, and High CI conditions. Panel <bold>(f)</bold> compares the normalized values of <inline-formula><mml:math id="M78" display="inline"><mml:mi mathvariant="normal">DEM</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M79" display="inline"><mml:mi mathvariant="normal">DEMstd</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M80" display="inline"><mml:mi mathvariant="normal">LAI</mml:mi></mml:math></inline-formula>, <inline-formula><mml:math id="M81" display="inline"><mml:mi mathvariant="normal">PFTh</mml:mi></mml:math></inline-formula>, and CI for these three representative basins.</p></caption>
            <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f03.png"/>

          </fig>

      <p id="d2e1756">Consequently, we consider a model to be robust in representing SMR processes if it not only performs well in basins with low basin complexity, but also maintains its accuracy as basin complexity increases. Our analysis therefore examines model performance against each complexity factor individually, and then assesses model performance along the integrated CI gradient.</p>
</sec>
<sec id="Ch1.S2.SS2.SSS4">
  <label>2.2.4</label><title>Measuring robustness of models</title>
      <p id="d2e1767">As described above, a robust model should accurately simulate SMR across diverse land-surface conditions. In this study, robustness is assessed by examining how model bias changes along the basin complexity gradient. Specifically, we developed a Robustness Index (RI) from two complementary perspectives: (1) the overall magnitude of model bias across all basin-complexity levels, hereafter referred to as performance stability; and (2) the rate at which model bias changes as basin complexity increases, hereafter referred to as adaptability.</p>
      <p id="d2e1770">Before calculating RI, all basins were ranked according to their integrated CI and divided into ten equal-sized groups along the CI gradient. Each group therefore represents one basin complexity level, from the least complex to the most complex basins. For each group, the representative CI value was defined as the median CI of all basins within that group, and the model bias was summarized using the median bias across basins in the same group. This grouping strategy reduces the influence of uneven basin distributions along the CI gradient and allows model performance to be evaluated consistently across different environmental complexity levels.</p>
      <p id="d2e1773">First, to quantify the overall bias magnitude, we calculated the Stratified Mean Absolute Bias (SMAB). SMAB represents the average absolute bias across the full CI gradient and is used here as a measure of performance stability. Compared with a direct average over all basins, this stratified calculation avoids potential distortion caused by uneven basin distributions along the complexity gradient and provides a more balanced estimate of model performance across different basin complexity levels. SMAB is calculated as follows (Eq. <xref ref-type="disp-formula" rid="Ch1.E5"/>):

              <disp-formula id="Ch1.E5" content-type="numbered"><label>5</label><mml:math id="M82" display="block"><mml:mrow><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mn mathvariant="normal">1</mml:mn><mml:mrow><mml:msub><mml:mi mathvariant="normal">CI</mml:mi><mml:mi>n</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:munderover><mml:mo movablelimits="false">∑</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:munderover><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:mfenced close="|" open="|"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Bias</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced><mml:mo>+</mml:mo><mml:mfenced close="|" open="|"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Bias</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="normal">CI</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi mathvariant="normal">CI</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M83" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula> denotes the index of the complexity groups, ranging from 1 to <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mi>n</mml:mi><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M85" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">Bias</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> represents the median model bias within the <inline-formula><mml:math id="M86" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>th complexity group, and <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">CI</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> indicates the representative CI value of that group.</p>
      <p id="d2e1909">Second, to quantify the sensitivity of model performance to increasing basin complexity, we regressed the absolute bias against CI using simple linear regression. The regression slope <inline-formula><mml:math id="M88" display="inline"><mml:mi>S</mml:mi></mml:math></inline-formula> was adopted as a measure of adaptability. A larger positive slope indicates that model bias increases more rapidly with basin complexity, implying stronger performance degradation and weaker adaptability under complex land-surface conditions. The regression is expressed as follows (Eq. <xref ref-type="disp-formula" rid="Ch1.E6"/>):

              <disp-formula id="Ch1.E6" content-type="numbered"><label>6</label><mml:math id="M89" display="block"><mml:mrow><mml:mo>|</mml:mo><mml:msub><mml:mi mathvariant="normal">Bias</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>|</mml:mo><mml:mo>=</mml:mo><mml:mi>S</mml:mi><mml:mo>×</mml:mo><mml:msub><mml:mi mathvariant="normal">CI</mml:mi><mml:mi>i</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mi>b</mml:mi></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M90" display="inline"><mml:mi>S</mml:mi></mml:math></inline-formula> is the slope of the absolute bias along the CI gradient, and <inline-formula><mml:math id="M91" display="inline"><mml:mi>b</mml:mi></mml:math></inline-formula> is the intercept.</p>
      <p id="d2e1966">For comparison across all models, both SMAB and <inline-formula><mml:math id="M92" display="inline"><mml:mi>S</mml:mi></mml:math></inline-formula> were normalized to a <inline-formula><mml:math id="M93" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> scale. The two normalized metrics describe complementary aspects of model robustness. Specifically, <inline-formula><mml:math id="M94" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> reflects the overall magnitude of model bias across the full basin complexity gradient and therefore represents performance stability. <inline-formula><mml:math id="M95" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> reflects the rate of bias increase with basin complexity and therefore represents adaptability to complex basin conditions. A robust model should have both low <inline-formula><mml:math id="M96" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and low <inline-formula><mml:math id="M97" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, indicating low overall bias and weak performance degradation as basin complexity increases.</p>
      <p id="d2e2037">Together, these two components form a two-dimensional metric space, represented by the vector <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:msub><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula>. In this space, the origin point <inline-formula><mml:math id="M99" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> represents an ideal model with zero average bias and no degradation trend along the CI gradient. Following previous studies that use distance-based approaches to integrate multiple performance dimensions <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx35 bib1.bibx22" id="paren.48"/>, we adopted the Euclidean distance from this ideal point to define RI. A larger distance indicates a greater deviation from the ideal condition and therefore lower robustness. The final RI is calculated as one minus this distance (Eq. <xref ref-type="disp-formula" rid="Ch1.E7"/>):

              <disp-formula id="Ch1.E7" content-type="numbered"><label>7</label><mml:math id="M100" display="block"><mml:mrow><mml:mi mathvariant="normal">RI</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mrow><mml:msubsup><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:msqrt></mml:mrow></mml:math></disp-formula></p>
      <p id="d2e2114">A higher RI value indicates that a model is closer to the ideal origin point, reflecting both lower overall bias and stronger adaptability to increasing basin complexity. In this formulation, models with large overall bias are penalized through <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, whereas models whose errors increase rapidly along the CI gradient are penalized through <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Therefore, RI provides an integrated measure of model robustness by jointly accounting for performance stability and adaptability. We note that RI is used here primarily as a relative robustness metric for inter-model comparison, rather than as an absolute measure of model performance.</p>
      <p id="d2e2139">To clarify the interpretation of RI, we added a conceptual illustration in the Appendix Fig. <xref ref-type="fig" rid="FA2"/>. The schematic shows that a model with low SMAB and low slope corresponds to high robustness, whereas a model with either large overall bias or strong bias increase along the CI gradient receives a lower RI. Meanwhile, to further examine whether the RI-based model ranking is strongly affected by the specific aggregation method, we conducted a set of sensitivity experiments using alternative RI formulations (Table <xref ref-type="table" rid="TA1"/> and Fig. <xref ref-type="fig" rid="FA3"/>). These experiments included weighted Euclidean-distance formulations and weighted linear formulations with different relative weights assigned to <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The results show that most models exhibit limited changes in both RI magnitude and relative ranking across the sensitivity experiments. In particular, models with consistently high or low robustness remain generally stable, indicating that the identification of robust and non-robust models is not strongly affected by the specific weighting scheme. Based on these sensitivity results, and considering that Euclidean-distance-based metrics have been widely used in previous studies to integrate multiple performance dimensions <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx35 bib1.bibx22" id="paren.49"/>, we retained the equal-weight Euclidean-distance formulation as the baseline RI calculation. This formulation provides a simple and physically interpretable way to jointly represent performance stability and adaptability, while avoiding the need to impose subjective preference on either component. Therefore, RI is used here as a relative and integrated measure of model robustness for inter-model comparison.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Results</title>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>The performance of 15 models and datasets</title>
      <p id="d2e2192">We first present the distribution of model biases in three key SMR characteristics (<inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ) across multiple models and datasets (Fig. <xref ref-type="fig" rid="F4"/>). It is evident from Fig. <xref ref-type="fig" rid="F4"/> that most models and datasets exhibit notable biases in simulating SMR characteristics. Specifically, 10 out of 15 tend to underestimate <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, while CTQ is generally predicted earlier than observed (14 out of 15). Although a few models perform consistently well, considerable variability exists across metrics. The following sections provide a detailed assessment of <inline-formula><mml:math id="M109" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ, respectively. Fig. <xref ref-type="fig" rid="F4"/>a shows that the highest accuracy of <inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is achieved by discharge datasets (i.e., GRDR and GRADES), followed by GHMs (blue bars) and LSMs (yellow bars). GRDR and GRADES have 42.4 % and 40.1 % of basins, respectively, that lie within the <inline-formula><mml:math id="M112" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % threshold, and with a median bias of <inline-formula><mml:math id="M113" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>10.3 % and 0.7 %. Among models, WATERGAP2-2E performs the best (30.1 %, <inline-formula><mml:math id="M114" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>1.8 %), while VIC, MATSIRO, and MIROC-INTEG-LAND show the poorest performance, each with median biases exceeding <inline-formula><mml:math id="M115" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>50 %. These results highlight persistent underestimation in many physically based models, while underscoring the advantage of observation-constrained datasets. Details of each model's performance are shown in Appendix Fig. <xref ref-type="fig" rid="FA4"/>.</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e2315"><italic>Evaluation of key SMR characteristics (averaged over 1979–2019) simulated by 15 models/datasets across 1455 basins.</italic>
<bold>(a)</bold>–<bold>(c)</bold> present total runoff (<inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), peak flow (<inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), and centroid timing of runoff (CTQ) of the SMR, respectively. The black dashed line denotes zero bias, and the red shading denotes the acceptable ranges (<inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % for <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> d for CTQ). Model rankings are determined by the proportion of basins falling within these ranges (e.g., red triangle). Model categories (e.g., GHMs, LSMs, DGVM, and Datasets) and ISIMIP phases are distinguished by background color and hatch patterns.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f04.png"/>

        </fig>

      <p id="d2e2397">Figure <xref ref-type="fig" rid="F4"/>b presents results for <inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, which largely resemble <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> but exhibit generally worse performance. The same three models (i.e., GRADES, GRDR, and WATERGAP2-2E) again lead in performance (with 31.8 %, 31.1 %, and 27.8 % of basins, respectively, lying within the acceptable range). The weakest performers remain unchanged (VIC, MATSIRO, MIROC-INTEG-LAND), but with lower proportions in the acceptable range (11.2 %, 10.0 %, and 4.0 %). HYDROPY achieves 20.3 % basin with bias within <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % for <inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (median bias: 23.3 %), but only 17.0 % for <inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (55.5 %). On average, models tend to simulate <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> with higher fidelity than <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (22.2 % vs. 19.3 %), reflecting greater challenges in capturing peak flows, which are more sensitive to melt rate and timing.</p>
      <p id="d2e2480">Model performance for CTQ (Fig. <xref ref-type="fig" rid="F4"/>c) displays a distinct pattern. GRDR remains the best, with 55.2 % of basins within <inline-formula><mml:math id="M129" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> d and a median bias of <inline-formula><mml:math id="M130" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>3 d. However, rankings among models diverge substantially from those for <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Notably, VIC, MATSIRO, and JULES-W2 – previously among the least accurate for <inline-formula><mml:math id="M133" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M134" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> – rank among the top performers for CTQ. Conversely, models such as DBH, which ranked fourth for <inline-formula><mml:math id="M135" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, perform poorly for CTQ. These contrasts reflect fundamental differences in how models represent the timing versus magnitude of snowmelt, and emphasize the importance of snowpack energy balance and melt onset processes in simulating runoff timing.</p>
      <p id="d2e2558">In addition, we compared the performance of different model categories and ISIMIP phase (Fig. <xref ref-type="fig" rid="F5"/>). The results show that GRDR and GRADES consistently outperform models, underscoring the importance of observational constraints. Among models, GHMs generally outperform LSMs in simulating <inline-formula><mml:math id="M136" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M137" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, whereas LSMs show better performance for CTQ (Fig. <xref ref-type="fig" rid="F5"/>e).</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e2589"><italic>Evaluation of simulated runoff over 1979–2019 across 1455 catchments grouped by model category and ISIMIP phase.</italic> Panels <bold>(a)</bold>, <bold>(c)</bold>, <bold>(e)</bold> show the bias distributions grouped by model category, including global hydrological models (GHMs), land surface models (LSMs), observation-constrained runoff datasets (Dataset), and the ensemble mean of all individual models (ENS_MEAN). Panels <bold>(b)</bold>, <bold>(d)</bold>, <bold>(f)</bold> show the corresponding results grouped by ISIMIP phase. Rows represent the three snowmelt runoff characteristics: total runoff (<inline-formula><mml:math id="M138" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>; <bold>a</bold>–<bold>b</bold>), peak flow (<inline-formula><mml:math id="M139" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>; <bold>c</bold>–<bold>d</bold>), and centroid timing of runoff (CTQ; <bold>e</bold>–<bold>f</bold>). The black dashed line indicates zero bias, and the gray dashed lines denote the acceptable ranges (<inline-formula><mml:math id="M140" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % for <inline-formula><mml:math id="M141" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M142" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and <inline-formula><mml:math id="M143" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> d for CTQ). Groups are ordered according to the proportion of basins falling within these acceptable ranges, with category-level rankings calculated from the average performance of individual models within each group.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f05.png"/>

        </fig>

      <p id="d2e2703">This contrast likely reflects the distinct physical controls underlying the three runoff characteristics. <inline-formula><mml:math id="M144" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M145" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are more closely related to runoff magnitude and water-balance closure during the snowmelt period, and therefore may benefit from runoff-generation parameterizations and streamflow-oriented calibration strategies that are commonly emphasized in GHMs. In contrast, CTQ mainly reflects the timing and rate of snowmelt release, which are more directly controlled by energy-exchange processes and snowpack evolution. This interpretation is also consistent with the process complexity analysis in Part 2 <xref ref-type="bibr" rid="bib1.bibx25" id="paren.50"/>, which shows that LSMs generally include more detailed energy-related process representations, such as energy-balance snowmelt schemes, canopy radiative transfer parameterizations. These features are particularly important for capturing melt timing and may partly explain why LSMs show relative advantages in CTQ simulation. Notably, two of the three models for CTQ are LSMs (Fig. <xref ref-type="fig" rid="F4"/>c), further suggesting that physically based energy-related process representations can improve the simulation of runoff timing. However, the differences between GHMs and LSMs should not be attributed solely to model category or process representation. Calibration status, forcing datasets, model resolution, and implementation strategies may also contribute to the observed patterns. For example, several GHMs in our ensemble have explicit calibration information, whereas most LSMs are grouped as uncalibrated models, which may partly contribute to the stronger GHM performance for <inline-formula><mml:math id="M146" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M147" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. In addition, models from ISIMIP3a outperform those from ISIMIP2a (Fig. <xref ref-type="fig" rid="F5"/>). This improvement may reflect a combination of updated model structures, differences in forcing datasets, and model-specific calibration or implementation strategies, rather than model generation alone. Therefore, the category-level differences reported here should be interpreted as the combined effect of runoff-characteristic sensitivities, process representation, calibration status, and forcing data differences.</p>
      <p id="d2e2758">We further evaluate the spatial performance of the models (Fig. <xref ref-type="fig" rid="F6"/>), which shows the overall biases, inter-model differences, and the best-performing model at each site.  Figure <xref ref-type="fig" rid="F6"/>a–c highlights the spatial patterns of ensemble mean simulation bias. Widespread underestimation and earlier runoff timing are observed across several regions, including the western coastal and mountainous areas of the United States, western and northern Europe, northeastern China, and Japan. In more than 50 % of the basins within these regions, <inline-formula><mml:math id="M148" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M149" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> exhibit negative biases ranging from 0 % to <inline-formula><mml:math id="M150" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>60 %, while CTQ occurs up to 15 d earlier than observed. Meanwhile, significant overestimations are found in the central United States and northern China. Across all basins, the proportion meeting the predefined performance thresholds, i.e., bias within <inline-formula><mml:math id="M151" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % for <inline-formula><mml:math id="M152" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M153" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and within <inline-formula><mml:math id="M154" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>5 d for CTQ, is 27.9 %, 29.4 %, and 45.0 %, respectively. These basins are mainly located along the eastern coast of the United States. Detailed spatial maps of each model's performance are shown in Appendix Figs. <xref ref-type="fig" rid="FA5"/>–<xref ref-type="fig" rid="FA7"/>.</p>
      <p id="d2e2839">We also assess inter-model consistency in Fig. <xref ref-type="fig" rid="F6"/>d–f, which shows the largest inter-model discrepancies occur in the central United States, northern Europe, and northeastern China, where CV for <inline-formula><mml:math id="M155" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M156" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> exceeds 0.8 and the CTQ range exceeds 30 d. Basins with relatively low variability, defined as CV <inline-formula><mml:math id="M157" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 0.4 or CTQ range <inline-formula><mml:math id="M158" display="inline"><mml:mo>&lt;</mml:mo></mml:math></inline-formula> 20 d, occupy 20.7 %, 15.2 %, and 24.6 % of all basins for <inline-formula><mml:math id="M159" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M160" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ, respectively. These basins are primarily located in mid- to low-latitude regions of the United States and western Europe.</p>
      <p id="d2e2903">Figure <xref ref-type="fig" rid="F6"/>g–i identifies the best-performing model in each basin, with the top two models highlighted in pie charts to examine whether any model demonstrates broad applicability. The results suggest that, no single model or dataset consistently outperforms others across all basins. For the SMR magnitude (i.e., <inline-formula><mml:math id="M161" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M162" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>), GRDR and GRADES emerge as the best ones, each being the best-performing model in over 10 % of basins. The remaining 70 % of basins are distributed among several other models, each contributing approximately 5 % on average. For the SMR timing (i.e., CTQ), DBH and PCR-GLOBWB are the most frequently selected, accounting for 14.1 % and 13.3 % of basins, respectively. These findings underscore the inconsistency in optimal model selection across different runoff characteristics.</p>
      <p id="d2e2930">Overall, basins exhibiting larger simulation biases tend to also display greater inter-model variability, indicating the persistent SMR challenges within these regions. This may be attributed to: (1) the complex land conditions, which makes it difficult for models to accurately represent relevant physical processes, and (2) differing model complexities, where simpler and more complex models diverge more significantly under such land conditions. These jointly determine the increased biases and reduced consistency there. Furthermore, model performance varies across basins–strong performance in one basin does not guarantee similar performance elsewhere. This highlights that no single model is universally applicable, and that model selection should consider both complexity of the basin and the model's ability to represent key physical processes.</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e2935"><italic>Bias, variability, and the best model for simulated SMR characteristics across different basins.</italic>
<bold>(a)</bold>–<bold>(c)</bold> show the percent bias of <inline-formula><mml:math id="M163" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (%), percent bias of <inline-formula><mml:math id="M164" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (%), and bias of CTQ (days), respectively, in the simulated SMR. <bold>(d)</bold>–<bold>(f)</bold> show inter-model variability represented by the coefficient of variation (CV) for <inline-formula><mml:math id="M165" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M166" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and inter-model range of CTQ (days). <bold>(g)</bold>–<bold>(i)</bold> show the model with the smallest bias in each basin for <inline-formula><mml:math id="M167" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M168" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ, respectively. The pie charts provide a summary of these patterns, with bold black outlines highlighting the proportion within <inline-formula><mml:math id="M169" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> % bias in <bold>(a)</bold>–<bold>(c)</bold>, the two lowest variability levels in <bold>(d)</bold>–<bold>(f)</bold>, and the top two models with the highest proportions in <bold>(g)</bold>–<bold>(i)</bold>.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f06.jpg"/>

        </fig>

</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>Impacts of land surface complexity on model performance</title>
      <p id="d2e3069">To understand how land surface complexity influences model performance, we next analyze the results to identify general patterns, dissect the effects of different complexity sources, and compare the behaviors of models with varying structures. The most pervasive pattern is that model performance deteriorates as basin complexity increases (Fig. <xref ref-type="fig" rid="F7"/>a–c). For the majority of models, biases in simulating <inline-formula><mml:math id="M170" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M171" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ grow as the land surface becomes more complex, confirming the hypothesis that many models have a limited ability to capture intricate land surface processes in such environments. Among the three characteristics, the bias in CTQ increases most markedly with basin complexity, followed by <inline-formula><mml:math id="M172" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M173" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Moreover, this fragility in CTQ is particularly evident beyond a certain complexity threshold (Fig. <xref ref-type="fig" rid="F7"/>c), where performance deteriorates sharply for nearly all models, suggesting a potential breakdown in their ability to handle compounded, non-linear landscape effects. For <inline-formula><mml:math id="M174" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M175" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, two distinct groups of models can be identified: those with consistently low bias (e.g., GRDR and GRADES) and those with consistently high bias (e.g., VIC). Notably, GRADES is based on VIC runoff simulations constrained by observations, while GRDR further assimilates remotely sensed river width information. The result provides an initial assessment of the gain from data assimilation as a strategy for correcting inherent model. Detailed relationships between each factor and the model performance can be found in the Appendix Figs. <xref ref-type="fig" rid="FA8"/>–<xref ref-type="fig" rid="FA10"/>.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e3149"><italic>Influence of basin complexity on model simulations of key runoff characteristics.</italic>
<bold>(a)</bold>–<bold>(c)</bold> show the variation in model performance as basin complexity increases, and <bold>(d)</bold>–<bold>(f)</bold> represent the Spearman correlation between model performance and individual basin complexity factors. A positive correlation indicates that model bias increases with increasing basin complexity, whereas a negative correlation indicates that model bias decreases as basin complexity increases. Shading indicates the basin complexity factor with the highest absolute Spearman correlation among the four factors.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f07.png"/>

        </fig>

      <p id="d2e3172">A deep dive into specific model comparisons offers interesting insights into how different model structures handle complexity. A noteworthy and counterintuitive observation is the improved performance of certain physically complex models in more challenging land surface. For example, models like JULES-W2 and PCR-GLOBWB exhibit an unexpected trend where their simulation bias for <inline-formula><mml:math id="M176" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> decreases as basin complexity increases (Fig. <xref ref-type="fig" rid="F7"/>a). This suggests that the advanced process representations within these models, which might be less critical in simple, homogeneous basins, become advantageous in complex terrain. For instance, JULES-W2’s sophisticated schemes for canopy interception, sublimation, and radiation transfer are explicitly designed to handle the fine-scale variability introduced by complex topography conditions (Fig. <xref ref-type="fig" rid="F7"/>d). However, we would like to note that while the result points to the potential of sophisticated physics to enhance model robustness, further diagnostic analysis is essential to confirm whether this improved performance reflects a genuinely better process representation.</p>
      <p id="d2e3191">The analysis also reveals significant structural trade-offs in how different models respond to complexity, particularly when comparing the simulation of <inline-formula><mml:math id="M177" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M178" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. This is exemplified by the opposing responses of DBH and LPJML: in complex basins, DBH improves its <inline-formula><mml:math id="M179" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (Fig. <xref ref-type="fig" rid="F7"/>b) simulation while degrading its <inline-formula><mml:math id="M180" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (Fig. <xref ref-type="fig" rid="F7"/>a) simulation, whereas LPJML exhibits the inverse pattern. These contrasting responses may reflect differences in how the two models simulate runoff generation for different flow characteristics. A model like DBH may have a structure that better represents the fast, event-based runoff pathways critical for capturing flood peaks, while LPJML might be better structured to simulate the slow, integrated processes that determine the long-term water balance. These findings imply that no single model structure currently excels at both functions in complex environments, highlighting the need to select models based on the specific scientific question at hand.</p>
      <p id="d2e3243">When the overall complexity is decomposed into its constituent components, a clear hierarchy of influence emerges (Fig. <xref ref-type="fig" rid="F7"/>d–f). The bias in simulated <inline-formula><mml:math id="M181" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is primarily driven by variations in DEM and PFTh (Fig. <xref ref-type="fig" rid="F7"/>d). This likely reflects the strong link between DEM and processes such as rainfall–snowfall partitioning and sublimation, while PFTh affects the accuracy of interception representation. These processes jointly determine the accuracy of <inline-formula><mml:math id="M182" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> simulation. In comparison, the bias in simulated <inline-formula><mml:math id="M183" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is more strongly influenced by vegetation cover (e.g., LAI) and vegetation type diversity (e.g., PFTh) (Fig. <xref ref-type="fig" rid="F7"/>e). Both LAI and PFTh affect the amount of energy and water reaching the ground surface; during the snowmelt period, the extent to which liquid precipitation interception is properly represented can substantially impact <inline-formula><mml:math id="M184" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. The timing metric CTQ is mainly controlled by topographic factors (Fig. <xref ref-type="fig" rid="F7"/>f), particularly DEM and its variability (DEMstd), likely because CTQ is more sensitive to energy-related processes, and both DEM and DEMstd are key determinants of surface energy balance. Overall, topography remains the dominant driver of model bias, underscoring its critical role in shaping runoff dynamics, while vegetation exerts a comparatively secondary influence. Their combined interactions ultimately govern the magnitude and structure of the overall model bias.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Assessment of Model Robustness</title>
      <p id="d2e3307">We further evaluate the aforementioned model's robustness by analyzing performance across the defined complexity groups (Fig. <xref ref-type="fig" rid="F8"/>). By focusing on two aspects of our robustness metric, we aim to identify significant performance differences among the models or model groups and explore the underlying reasons for these variations. Models with lower stability and lower adaptability are considered higher robust in reproducing runoff characteristics (Fig. <xref ref-type="fig" rid="F8"/>a).</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e3316"><italic>Assessment of model robustness in key SMR characteristics.</italic>
<bold>(a)</bold> Schematic illustration. <bold>(b–d)</bold> Results for <inline-formula><mml:math id="M185" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M186" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ, respectively. Circles denote global hydrological models (GHMs), squares denote land surface models (LSMs), triangles denote dynamic global vegetation models (DGVMs), and diamonds denote data products. <inline-formula><mml:math id="M187" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes the normalized stratified mean absolute bias, and <inline-formula><mml:math id="M188" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> denotes the normalized regression slope of absolute bias. Detailed calculation procedures are provided in Sect. 2.2.4.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f08.jpg"/>

        </fig>

      <fig id="F9" specific-use="star"><label>Figure 9</label><caption><p id="d2e3380"><italic>Assessment of model robustness in key SMR characteristics grouped by model categories and ISIMIP phase.</italic>
<bold>(a)</bold>–<bold>(c)</bold> show results for <inline-formula><mml:math id="M189" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M190" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ, respectively. The orange line denotes the median, the black triangle indicates the mean, and the classifications of model types and ISIMIP phases follow Sect. 2.1.1.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f09.png"/>

        </fig>

      <p id="d2e3420">Overall, the models with the highest robustness for <inline-formula><mml:math id="M191" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are GRDR, GRADES, and PCR-GLOBWB (Fig. <xref ref-type="fig" rid="F8"/>b). These models exhibit both low average bias and minimal performance degradation as land surface complexity increases. In contrast, models like CLM40, MATSIRO, and MIROC-INTEG-LAND rank the lowest, primarily due to their large simulation biases. The robustness rankings for <inline-formula><mml:math id="M192" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are generally consistent, with GRDR, GRADES again performing well (Fig. <xref ref-type="fig" rid="F8"/>c). However, a key difference is that several models, such as CWATM and HYDROPY, exhibit significant improvements in simulating <inline-formula><mml:math id="M193" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> in more complex basins, hinting at structural differences in how they handle event-based runoff.</p>
      <p id="d2e3460">The results for CTQ reveal a more critical pattern (Fig. <xref ref-type="fig" rid="F8"/>d). While the overall bias is comparable across many models, the performance of nearly all models degrades sharply with increasing complexity. This indicates that runoff timing is the characteristic most sensitive to the challenges posed by complex land surfaces, as indicated in Sect. 3.2. This S<sub>norm</sub> arises largely because timing is influenced by coupled processes that are difficult to simulate accurately in complex environments. For example, topography modulates surface radiation and alters runoff convergence, while dense and diverse vegetation complicates canopy interception and energy transfer. The inability of current models to resolve these intricate interactions leads to the sharp decline in performance for accurately predicting the timing of the snowmelt.</p>
      <p id="d2e3474">A deep dive into the two components of the robustness score provides more insights in model behavior. For <inline-formula><mml:math id="M195" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, some models present a clear conflict between these two metrics (Fig. <xref ref-type="fig" rid="F8"/>b). Models like LPJML and CLM40 may appear accurate in simple basins, but their performance degrades rapidly in more complex environments (Fig. <xref ref-type="fig" rid="F7"/>a), resulting in a higher slope and consequently greater instability in performance. Conversely, although PCR-GLOBWB exhibits a relatively large overall bias, its performance improves with increasing basin complexity (Fig. <xref ref-type="fig" rid="F7"/>a), thereby achieving a high robustness ranking. This seemingly counterintuitive result may reflect a trade-off inherent in physically complex models, as their sophisticated process representations provide advantages for simulating the intricate dynamics in challenging basins but may introduce unnecessary structural uncertainty in simpler environments at the same time. Such a pattern is even more pronounced for <inline-formula><mml:math id="M196" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (see Fig. <xref ref-type="fig" rid="F8"/>c), where a larger set of models, including DBH and HYDROPY, exhibit a low slope. This suggests their internal structures may be better suited to capturing the dynamics of peak discharge in highly complex terrain conditions.</p>
      <p id="d2e3508">We further aggregate the results by model group to identify systematic differences in performance (Fig. <xref ref-type="fig" rid="F9"/>). Overall, ISIMIP3a outperforms ISIMIP2a in simulating <inline-formula><mml:math id="M197" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M198" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Moreover, observation-constrained datasets (GRADES and GRDR) exhibit the highest robustness, followed by GHMs and LSMs. These GHMs advantage is particularly pronounced for <inline-formula><mml:math id="M199" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M200" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (Fig. <xref ref-type="fig" rid="F9"/>a–b). For CTQ (Fig. <xref ref-type="fig" rid="F9"/>c), however, LSMs exhibit a clear relative improvement compared with their performance for <inline-formula><mml:math id="M201" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M202" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, with MIROC-INTEG-LAND notably rising from the lowest tier to a leading position (Fig. <xref ref-type="fig" rid="F8"/>d). This improvement likely reflects the more advanced treatment of radiative transfer and surface energy balance in LSMs, which plays a greater role in capturing melt timing than in reproducing total runoff or peak flow.</p>
</sec>
</sec>
<sec id="Ch1.S4" sec-type="conclusions">
  <label>4</label><title>Conclusions and Discussions</title>
      <p id="d2e3595">This study provides a systematic large-sample evaluation of the ability of 15 hydrological models and runoff products to capture key runoff characteristics during the snowmelt period across 1455 basins worldwide. In addition to conventional performance assessment, we use a RI to quantify how model skill changes along basin complexity gradients. This analysis complements existing model evaluation approaches by integrating established SMR characteristics with basin complexity information, thereby providing additional insights into the stability and adaptability of model performance under diverse environmental conditions. Our primary findings are summarized as follows: <list list-type="bullet"><list-item>
      <p id="d2e3600">Most models exhibit systematic biases in simulating key runoff characteristics during the snowmelt period. In particular, they tend to underestimate <inline-formula><mml:math id="M203" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M204" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> while predicting CTQ too early. These biases are especially pronounced in regions such as the western United States, northern Europe, and northeastern China, where both the magnitude of deviations and the inter-model discrepancies are substantial. GRDR and GRADES generally outperform most process-based models in reproducing SMR characteristics, highlighting the value of observational constraints. Among them, GHMs generally perform better than LSMs in simulating runoff magnitude characteristics, including <inline-formula><mml:math id="M205" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M206" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, whereas LSMs show a clear advantage over GHMs in simulating CTQ. This contrast suggests that GHMs are more effective in reproducing snowmelt runoff magnitude, while LSMs are better suited to capturing energy-controlled melt timing. The ISIMIP3a models also show consistently higher accuracy compared to ISIMIP2a, and ensemble means typically provide more robust results than individual models.</p></list-item><list-item>
      <p id="d2e3648">Model biases are substantially increased under high basin complexity, with stronger underestimation of <inline-formula><mml:math id="M207" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M208" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and earlier estimation of CTQ. The underestimation is generally more severe for <inline-formula><mml:math id="M209" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> than for <inline-formula><mml:math id="M210" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>. Notably, model skill in simulating CTQ declines sharply as basin complexity increases, highlighting the limited capacity of current models to capture complex snow–vegetation–topography interactions under highly heterogeneous conditions. Models show divergent responses to increasing basin complexity: GRDR and GRADES remain relatively robust, some models (e.g., JULES-W2, PCR-GLOBWB) show partial adaptability, while many others exhibit increasing biases.</p></list-item><list-item>
      <p id="d2e3696">By applying the newly developed robustness metric, we find that GRDR and GRADES consistently rank at the top, with PCR-GLOBWB, JULES-W2, and DBH also performing well. By contrast, CLM40, MATSIRO, and MIROC-INTEG-LAND show low robustness for <inline-formula><mml:math id="M211" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M212" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> but improve markedly in CTQ. Overall, GHMs demonstrate higher robustness in simulating runoff magnitude characteristics, particularly <inline-formula><mml:math id="M213" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M214" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, whereas LSMs show higher robustness in simulating CTQ. This contrast likely reflects fundamental differences in model structure: GHMs are generally more effective in maintaining stable water-balance and runoff-generation performance, while LSMs benefit from more detailed energy-balance and snowpack representations that improve the robustness of melt-timing simulation under heterogeneous conditions.</p></list-item></list> These findings advance the understanding of SMR modeling under complex land surface conditions, establishing a benchmark framework for future model development. Several limitations are worthy of further discussion for future research. <list list-type="bullet"><list-item>
      <p id="d2e3746">First, it is noteworthy that ISIMIP2a and ISIMIP3a exhibit differences beyond their mechanistic descriptions of the snow accumulation and snowmelt processes. These variations include differences in forcing, simulation scenarios, and whether model calibration is applied. In our data selection process, we have minimized these discrepancies to enhance direct comparisons that emphasize process differences. However, further scrutiny may be necessary to fully delineate these differences.</p></list-item><list-item>
      <p id="d2e3750">Second, the evaluation metrics of models could be further improved. While this study primarily focused on biases and robustness in simulating key SMR characteristics, it did not assess the models’ capability in reproducing the temporal dynamics of SMR. Future research should therefore focus on evaluating how well models capture the temporal dynamics of SMR.</p></list-item><list-item>
      <p id="d2e3754">Third, the relationship between model performance and the completeness of physical process representation requires further investigation. Current analyses are largely conducted at the level of model categories or intercomparison projects, providing only a broad perspective on the link between model structure and performance. Future work should therefore focus on developing a unified framework to quantify model process complexity, which would allow a more rigorous evaluation of how physical process differences translate into performance disparities.</p></list-item></list></p>
</sec>

      
      </body>
    <back><app-group>

<app id="App1.Ch1.S1">
  <label>Appendix A</label><title/>

<table-wrap id="TA1"><label>Table A1</label><caption><p id="d2e3772"><italic>Sensitivity experiments used to test the robustness of the RI formulation.</italic> Formula A: <inline-formula><mml:math id="M215" display="inline"><mml:mrow><mml:mi mathvariant="normal">RI</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mrow><mml:mi>w</mml:mi><mml:mo>⋅</mml:mo><mml:msubsup><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi>w</mml:mi><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:msubsup><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi><mml:mn mathvariant="normal">2</mml:mn></mml:msubsup></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula>. Formula B: <inline-formula><mml:math id="M216" display="inline"><mml:mrow><mml:mi mathvariant="normal">RI</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mfenced open="(" close=")"><mml:mrow><mml:mi>w</mml:mi><mml:mo>⋅</mml:mo><mml:msub><mml:mi mathvariant="normal">SMAB</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mi>w</mml:mi><mml:mo>)</mml:mo><mml:mo>⋅</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mi mathvariant="normal">norm</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula>.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Experiment</oasis:entry>
         <oasis:entry colname="col2">Formulation</oasis:entry>
         <oasis:entry colname="col3">Weight setting</oasis:entry>
         <oasis:entry colname="col4">Purpose</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">S0</oasis:entry>
         <oasis:entry colname="col2">A</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M217" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Baseline RI formulation.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S1-1</oasis:entry>
         <oasis:entry colname="col2">A</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M218" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Euclidean aggregation; adaptability emphasized.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S1-2</oasis:entry>
         <oasis:entry colname="col2">A</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M219" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.4</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Euclidean aggregation; adaptability slightly emphasized.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S1-3</oasis:entry>
         <oasis:entry colname="col2">A</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M220" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.6</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Euclidean aggregation; stability slightly emphasized.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S1-4</oasis:entry>
         <oasis:entry colname="col2">A</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M221" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Euclidean aggregation; stability emphasized.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S2-1</oasis:entry>
         <oasis:entry colname="col2">B</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M222" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Linear aggregation; adaptability emphasized.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S2-2</oasis:entry>
         <oasis:entry colname="col2">B</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M223" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.4</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Linear aggregation; adaptability slightly emphasized.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S2-3</oasis:entry>
         <oasis:entry colname="col2">B</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M224" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.5</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Linear aggregation; equal weighting.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S2-4</oasis:entry>
         <oasis:entry colname="col2">B</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M225" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.6</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Linear aggregation; stability slightly emphasized.</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">S2-5</oasis:entry>
         <oasis:entry colname="col2">B</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M226" display="inline"><mml:mrow><mml:mi>w</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.7</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Linear aggregation; stability emphasized.</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

<fig id="FA1"><label>Figure A1</label><caption><p id="d2e4158"><italic>Pearson correlation matrix among the four components used to construct the basin complexity index.</italic> The value of each grid represents the Pearson correlation coefficients among teh DEM, DEMstd, LAI, and PFTh.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f10.png"/>

      </fig>

      <fig id="FA2"><label>Figure A2</label><caption><p id="d2e4174"><italic>Conceptual illustration of the Robustness Index (RI).</italic> The stratified mean absolute bias (SMAB) represents the overall magnitude of model bias across the full basin complexity index (CI) gradient, whereas the slope represents the rate at which model bias changes with increasing CI.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f11.png"/>

      </fig>

<fig id="FA3"><label>Figure A3</label><caption><p id="d2e4190"><italic>Sensitivity analysis of the Robustness Index (RI) formulation.</italic> Panels <bold>(a)</bold>–<bold>(c)</bold> present the sensitivity of model robustness rankings for <inline-formula><mml:math id="M227" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M228" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>, and CTQ, respectively. The sensitivity experiments include the baseline Euclidean-distance-based RI (S0), weighted Euclidean formulations (S1-1 to S1-4), and weighted linear formulations (S2-1 to S2-5). The color shading represents the RI value, and the number in each cell denotes the model rank under the corresponding experiment. Red numbers indicate rank changes relative to the baseline RI formulation.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f12.png"/>

      </fig>

<fig id="FA4"><label>Figure A4</label><caption><p id="d2e4234"><italic>Percentage of gauges (%) with well-simulated snowmelt runoff indices for each model.</italic>
<bold>(a)</bold> shows the percentage of gauges with good simulated CTQ (within <inline-formula><mml:math id="M229" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> d of the observed) as ranked by their performance. <bold>(b)</bold> and <bold>(c)</bold> show that with well-simulated <inline-formula><mml:math id="M230" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (PBias within <inline-formula><mml:math id="M231" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> %) and that with well-simulated <inline-formula><mml:math id="M232" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> (PBias within <inline-formula><mml:math id="M233" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> %) in the snowmelt period.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f13.png"/>

      </fig>

<fig id="FA5"><label>Figure A5</label><caption><p id="d2e4313"><italic>The percentage bias of total runoff (</italic><inline-formula><mml:math id="M234" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="italic">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula><italic>) during the snowmelt period (unit: %).</italic> Colors indicate the degree of bias compared to observations: red represents underestimation, blue indicates overestimation, and yellow signifies well-matched total runoff. The point size is adjusted based on station density for better visualization and does not have physical significance. Colored boxes around model/dataset names denote three categories of data.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f14.jpg"/>

      </fig>

<fig id="FA6"><label>Figure A6</label><caption><p id="d2e4341"><italic>The percentage bias of peak flow (</italic><inline-formula><mml:math id="M235" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="italic">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula><italic>) during the snowmelt period (unit: %).</italic> Colors indicate the degree of bias compared to observations: red represents underestimation, blue indicates overestimation, and yellow signifies well-matched total runoff. The point size is adjusted based on station density for better visualization and does not have physical significance. Colored boxes around model/dataset names denote three categories of analysis data.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f15.jpg"/>

      </fig>

<fig id="FA7"><label>Figure A7</label><caption><p id="d2e4369"><italic>The bias in snowmelt timing, represented by the centroid time of runoff (CTQ) during the snowmelt period (unit: day)</italic> Colors indicate the extent of bias compared to observations: red represents earlier snowmelt, blue indicates later snowmelt, and yellow signifies well-matched snowmelt timing. The point size is adjusted based on station density for better visualization and does not have physical significance. Colored boxes around model/dataset names denote three categories of analysis data.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f16.jpg"/>

      </fig>

<fig id="FA8"><label>Figure A8</label><caption><p id="d2e4386"><italic>The bias in total runoff (</italic><inline-formula><mml:math id="M236" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="italic">sum</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula><italic>) in the snowmelt period as a function of basin complexity factors.</italic> The value of each grid represents the median bias (color shading) within the corresponding range of DEM, DEMstd, LAI, and PFTh. The figures are ordered from top to bottom based on the proportion of biases within <inline-formula><mml:math id="M237" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> %.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f17.jpg"/>

      </fig>

<fig id="FA9"><label>Figure A9</label><caption><p id="d2e4424"><italic>The bias in peak flow (</italic><inline-formula><mml:math id="M238" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="italic">max</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula><italic>) in the snowmelt period as a function of basin complexity factors.</italic> The value of each grid represents the median bias (color shading) within the corresponding range of DEM, DEMstd, LAI, and PFTh. The figures are ordered from top to bottom based on the proportion of biases within <inline-formula><mml:math id="M239" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">20</mml:mn></mml:mrow></mml:math></inline-formula> %.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f18.jpg"/>

      </fig>

<fig id="FA10"><label>Figure A10</label><caption><p id="d2e4462"><italic>The bias in centroid timing of runoff (CTQ) during in the snowmelt period as a function of basin complexity factors.</italic> The value of each grid represents the median bias (color shading) within the corresponding range of DEM, DEMstd, LAI, and PFTh. The figures are ordered from top to bottom based on the proportion of biases within <inline-formula><mml:math id="M240" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">5</mml:mn></mml:mrow></mml:math></inline-formula> d.</p></caption>
        
        <graphic xlink:href="https://hess.copernicus.org/articles/30/5343/2026/hess-30-5343-2026-f19.jpg"/>

      </fig>


</app>
  </app-group><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d2e4491">All data used in this study are available from public repositories: (a) ISIMIP model outputs from <uri>https://data.isimip.org/</uri>; (b) GRADES <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx30" id="paren.51"/> from <ext-link xlink:href="https://doi.org/10.11888/Terre.tpdc.272898" ext-link-type="DOI">10.11888/Terre.tpdc.272898</ext-link>; (c) GRDR <xref ref-type="bibr" rid="bib1.bibx10 bib1.bibx11" id="paren.52"/> from <ext-link xlink:href="https://doi.org/10.5281/zenodo.13951712" ext-link-type="DOI">10.5281/zenodo.13951712</ext-link>; (d) GSHA <xref ref-type="bibr" rid="bib1.bibx50 bib1.bibx49" id="paren.53"/> from <ext-link xlink:href="https://doi.org/10.5281/zenodo.10433905" ext-link-type="DOI">10.5281/zenodo.10433905</ext-link>. The code used in this study is available from the corresponding author upon reasonable request.</p>
  </notes><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e4519">Conceptualization: PL, XL, HL. Investigation: XL, HL, PL. Data curation: XL, HL. Funding acquisition: PL. Investigation: XL, HL, PL. Methodology: XL, PL, HL, KZ. Visualization: XL, HL, KZ. Writing (initial): XL, HL, PL. Writing (review and editing): XL, HL, PL, KZ.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e4525">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e4531">Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e4537">This study was supported by the National Key Research and Development Program of China (2022YFF0801303), the Beijing Nova Program (20230484302), the Beijing Nova Interdisciplinary Program (20240484647), the National Natural Science Foundation of China (42371481), and the Yunnan Provincial Science and Technology Project at Southwest United Graduate School (202302AO370012). The authors acknowledge valuable feedback from ISIMIP modelers Drs. Yusuke Satoh and Emmanouil Grillakis. We also thank Dr. Dashan Wang for insightful discussions related to this project.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e4542">This research has been supported by the National Key  Research and Development Program of China (grant no. 2022YFF0801303), the Beijing Nova Program (grant no. 20230484302), the Beijing Nova Interdisciplinary Program (grant no. 20240484647), the National Natural Science Foundation of China (grant no. 42371481), and the Yunnan Provincial Science and Technology Project at Southwest United Graduate School  (grant no. 202302AO370012).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e4548">This paper was edited by Xing Yuan and reviewed by two anonymous referees.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Andreadis et al.(2013)Andreadis, Schumann, and Pavelsky</label><mixed-citation>Andreadis, K. M., Schumann, G. J.-P., and Pavelsky, T.: A Simple Global River Bankfull Width and Depth Database: Data and Analysis Note, Water Resour. Res., 49, 7164–7168, <ext-link xlink:href="https://doi.org/10.1002/wrcr.20440" ext-link-type="DOI">10.1002/wrcr.20440</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Beck et al.(2017)Beck, van Dijk, de Roo, Dutra, Fink, Orth, and Schellekens</label><mixed-citation>Beck, H. E., van Dijk, A. I. J. M., de Roo, A., Dutra, E., Fink, G., Orth, R., and Schellekens, J.: Global evaluation of runoff from 10 state-of-the-art hydrological models, Hydrol. Earth Syst. Sci., 21, 2881–2903, <ext-link xlink:href="https://doi.org/10.5194/hess-21-2881-2017" ext-link-type="DOI">10.5194/hess-21-2881-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Beck et al.(2019)Beck, Wood, Pan, Fisher, Miralles, van Dijk, McVicar, and Adler</label><mixed-citation>Beck, H. E., Wood, E. F., Pan, M., Fisher, C. K., Miralles, D. G., van Dijk, A. I. J. M., McVicar, T. R., and Adler, R. F.: MSWEP V2 Global 3-Hourly 0.1<sup>∘</sup> Precipitation: Methodology and Quantitative Assessment, B. Am. Meteorol. Soc, 100, 473–500, <ext-link xlink:href="https://doi.org/10.1175/BAMS-D-17-0138.1" ext-link-type="DOI">10.1175/BAMS-D-17-0138.1</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Best et al.(2011)Best, Pryor, Clark, Rooney, Essery, Ménard, Edwards, Hendry, Porson, Gedney, Mercado, Sitch, Blyth, Boucher, Cox, Grimmond, and Harding</label><mixed-citation>Best, M. J., Pryor, M., Clark, D. B., Rooney, G. G., Essery, R. L. H., Ménard, C. B., Edwards, J. M., Hendry, M. A., Porson, A., Gedney, N., Mercado, L. M., Sitch, S., Blyth, E., Boucher, O., Cox, P. M., Grimmond, C. S. B., and Harding, R. J.: The Joint UK Land Environment Simulator (JULES), model description – Part 1: Energy and water fluxes, Geosci. Model Dev., 4, 677–699, <ext-link xlink:href="https://doi.org/10.5194/gmd-4-677-2011" ext-link-type="DOI">10.5194/gmd-4-677-2011</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Burek et al.(2020)Burek, Satoh, Kahil, Tang, Greve, Smilovic, Guillaumot, Zhao, and Wada</label><mixed-citation>Burek, P., Satoh, Y., Kahil, T., Tang, T., Greve, P., Smilovic, M., Guillaumot, L., Zhao, F., and Wada, Y.: Development of the Community Water Model (CWatM v1.04) – a high-resolution hydrological model for global and regional assessment of integrated water resources management, Geosci. Model Dev., 13, 3267–3298, <ext-link xlink:href="https://doi.org/10.5194/gmd-13-3267-2020" ext-link-type="DOI">10.5194/gmd-13-3267-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Chai et al.(2025)Chai, Miao, Gentine, Mudryk, Thackeray, Berghuijs, Wu, Fan, Slater, Sun, and Zwiers</label><mixed-citation>Chai, Y., Miao, C., Gentine, P., Mudryk, L., Thackeray, C. W., Berghuijs, W. R., Wu, Y., Fan, X., Slater, L., Sun, Q., and Zwiers, F.: Constrained Earth System Models Show a Stronger Reduction in Future Northern Hemisphere Snowmelt Water, Nat. Clim. Change, 15, <ext-link xlink:href="https://doi.org/10.1038/s41558-025-02308-y" ext-link-type="DOI">10.1038/s41558-025-02308-y</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Chen et al.(2021)Chen, Liu, Mao, Wang, Zeng, Chen, Wang, and Chen</label><mixed-citation>Chen, H., Liu, J., Mao, G., Wang, Z., Zeng, Z., Chen, A., Wang, K., and Chen, D.: Intercomparison of Ten ISI-MIP Models in Simulating Discharges along the Lancang-Mekong River Basin, Sci. Total Environ, 765, 144494, <ext-link xlink:href="https://doi.org/10.1016/j.scitotenv.2020.144494" ext-link-type="DOI">10.1016/j.scitotenv.2020.144494</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>David et al.(2011)David, Maidment, Niu, Yang, Habets, and Eijkhout</label><mixed-citation>David, C. H., Maidment, D. R., Niu, G.-Y., Yang, Z.-L., Habets, F., and Eijkhout, V.: River Network Routing on the NHDPlus Dataset, J. Hydrometeorol, 12, 913–934, <ext-link xlink:href="https://doi.org/10.1175/2011JHM1345.1" ext-link-type="DOI">10.1175/2011JHM1345.1</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Dudley et al.(2017)Dudley, Hodgkins, McHale, Kolian, and Renard</label><mixed-citation>Dudley, R., Hodgkins, G., McHale, M., Kolian, M., and Renard, B.: Trends in Snowmelt-Related Streamflow Timing in the Conterminous United States, J. Hydrol., 547, 208–221, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2017.01.051" ext-link-type="DOI">10.1016/j.jhydrol.2017.01.051</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Feng and Gleason(2024a)</label><mixed-citation>Feng, D. and Gleason, C. J.: More Flow Upstream and Less Flow Downstream: The Changing Form and Function of Global Rivers, Science, 386, 1305–1311, <ext-link xlink:href="https://doi.org/10.1126/science.adl5728" ext-link-type="DOI">10.1126/science.adl5728</ext-link>, 2024a.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Feng and Gleason(2024b)</label><mixed-citation>Feng, D. and Gleason, C.: Global River Discharge Reanalysis dataset (GRDR), Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.13951712" ext-link-type="DOI">10.5281/zenodo.13951712</ext-link>, 2024b.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Fenicia et al.(2014)Fenicia, Kavetski, Savenije, Clark, Schoups, Pfister, and Freer</label><mixed-citation>Fenicia, F., Kavetski, D., Savenije, H. H. G., Clark, M. P., Schoups, G., Pfister, L., and Freer, J.: Catchment Properties, Function, and Conceptual Model Representation: Is There a Correspondence?, Hydrol. Processes, 28, 2451–2467, <ext-link xlink:href="https://doi.org/10.1002/hyp.9726" ext-link-type="DOI">10.1002/hyp.9726</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Guimberteau et al.(2018)Guimberteau, Zhu, Maignan, Huang, Yue, Dantec-Nédélec, Ottlé, Jornet-Puig, Bastos, Laurent, Goll, Bowring, Chang, Guenet, Tifafi, Peng, Krinner, Ducharne, Wang, Wang, Wang, Wang, Yin, Lauerwald, Joetzjer, Qiu, Kim, and Ciais</label><mixed-citation>Guimberteau, M., Zhu, D., Maignan, F., Huang, Y., Yue, C., Dantec-Nédélec, S., Ottlé, C., Jornet-Puig, A., Bastos, A., Laurent, P., Goll, D., Bowring, S., Chang, J., Guenet, B., Tifafi, M., Peng, S., Krinner, G., Ducharne, A., Wang, F., Wang, T., Wang, X., Wang, Y., Yin, Z., Lauerwald, R., Joetzjer, E., Qiu, C., Kim, H., and Ciais, P.: ORCHIDEE-MICT (v8.4.1), a land surface model for the high latitudes: model description and validation, Geosci. Model Dev., 11, 121–163, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-121-2018" ext-link-type="DOI">10.5194/gmd-11-121-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Guo et al.(2024)Guo, Hou, Yang, and Mcvicar</label><mixed-citation>Guo, H., Hou, Y., Yang, Y., and Mcvicar, T. R.: Global Evaluation of Simulated High and Low Flows from 23 Macroscale Models, J. Hydrometeorol, 25, 425–443, <ext-link xlink:href="https://doi.org/10.1175/JHM-D-23-0176.1" ext-link-type="DOI">10.1175/JHM-D-23-0176.1</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Gupta et al.(2009)Gupta, Kling, Yilmaz, and Martinez</label><mixed-citation>Gupta, H. V., Kling, H., Yilmaz, K. K., and Martinez, G. F.: Decomposition of the Mean Squared Error and NSE Performance Criteria: Implications for Improving Hydrological Modelling, J. Hydrol., 377, 80–91, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2009.08.003" ext-link-type="DOI">10.1016/j.jhydrol.2009.08.003</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Haddeland et al.(2011)Haddeland, Clark, Franssen, Ludwig, Voß, Arnell, Bertrand, Best, Folwell, Gerten, Gomes, Gosling, Hagemann, Hanasaki, Harding, Heinke, Kabat, Koirala, Oki, Polcher, Stacke, Viterbo, Weedon, and Yeh</label><mixed-citation>Haddeland, I., Clark, D. B., Franssen, W., Ludwig, F., Voß, F., Arnell, N. W., Bertrand, N., Best, M., Folwell, S., Gerten, D., Gomes, S., Gosling, S. N., Hagemann, S., Hanasaki, N., Harding, R., Heinke, J., Kabat, P., Koirala, S., Oki, T., Polcher, J., Stacke, T., Viterbo, P., Weedon, G. P., and Yeh, P.: Multimodel Estimate of the Global Terrestrial Water Balance: Setup and First Results, J. Hydrometeorol, 12, 869–884, <ext-link xlink:href="https://doi.org/10.1175/2011JHM1324.1" ext-link-type="DOI">10.1175/2011JHM1324.1</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Han et al.(2024)Han, Liu, Woods, McVicar, Yang, Wang, Hou, Guo, Li, and Yang</label><mixed-citation>Han, J., Liu, Z., Woods, R., McVicar, T. R., Yang, D., Wang, T., Hou, Y., Guo, Y., Li, C., and Yang, Y.: Streamflow Seasonality in a Snow-Dwindling World, Nature, 629, 1075–1081, <ext-link xlink:href="https://doi.org/10.1038/s41586-024-07299-y" ext-link-type="DOI">10.1038/s41586-024-07299-y</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Hanasaki et al.(2008)Hanasaki, Kanae, Oki, Masuda, Motoya, Shirakawa, Shen, and Tanaka</label><mixed-citation>Hanasaki, N., Kanae, S., Oki, T., Masuda, K., Motoya, K., Shirakawa, N., Shen, Y., and Tanaka, K.: An integrated model for the assessment of global water resources – Part 1: Model description and input meteorological forcing, Hydrol. Earth Syst. Sci., 12, 1007–1025, <ext-link xlink:href="https://doi.org/10.5194/hess-12-1007-2008" ext-link-type="DOI">10.5194/hess-12-1007-2008</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Harper et al.(2023)Harper, Lamarche, Hartley, Peylin, Ottlé, Bastrikov, San Martín, Bohnenstengel, Kirches, Boettcher, Shevchuk, Brockmann, and Defourny</label><mixed-citation>Harper, K. L., Lamarche, C., Hartley, A., Peylin, P., Ottlé, C., Bastrikov, V., San Martín, R., Bohnenstengel, S. I., Kirches, G., Boettcher, M., Shevchuk, R., Brockmann, C., and Defourny, P.: A 29-year time series of annual 300 m resolution plant-functional-type maps for climate models, Earth Syst. Sci. Data, 15, 1465–1499, <ext-link xlink:href="https://doi.org/10.5194/essd-15-1465-2023" ext-link-type="DOI">10.5194/essd-15-1465-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Hersbach et al.(2020)Hersbach, Bell, Berrisford, Hirahara, Horányi, Muñoz-Sabater, Nicolas, Peubey, Radu, Schepers, Simmons, Soci, Abdalla, Abellan, Balsamo, Bechtold, Biavati, Bidlot, Bonavita, De Chiara, Dahlgren, Dee, Diamantakis, Dragani, Flemming, Forbes, Fuentes, Geer, Haimberger, Healy, Hogan, Hólm, Janisková, Keeley, Laloyaux, Lopez, Lupu, Radnoti, de Rosnay, Rozum, Vamborg, Villaume, and Thépaut</label><mixed-citation>Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A., Muñoz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D., Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P., Biavati, G., Bidlot, J., Bonavita, M., De Chiara, G., Dahlgren, P., Dee, D., Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer, A., Haimberger, L., Healy, S., Hogan, R. J., Hólm, E., Janisková, M., Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., de Rosnay, P., Rozum, I., Vamborg, F., Villaume, S., and Thépaut, J.-N.: The ERA5 Global Reanalysis, Q. J. Roy. Meteor. Soc., 146, 1999–2049, <ext-link xlink:href="https://doi.org/10.1002/qj.3803" ext-link-type="DOI">10.1002/qj.3803</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Hou et al.(2023)Hou, Guo, Yang, and Liu</label><mixed-citation>Hou, Y., Guo, H., Yang, Y., and Liu, W.: Global Evaluation of Runoff Simulation from Climate, Hydrological and Land Surface Models, Water Resour. Res., 59, e2021WR031817, <ext-link xlink:href="https://doi.org/10.1029/2021WR031817" ext-link-type="DOI">10.1029/2021WR031817</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Hu et al.(2022)Hu, Chen, Chen, Zhou, Peng, Li, and Sang</label><mixed-citation>Hu, Z., Chen, D., Chen, X., Zhou, Q., Peng, Y., Li, J., and Sang, Y.: CCHZ-DISO: A Timely New Assessment System for Data Quality or Model Performance From Da Dao Zhi Jian, Geophys. Res. Lett, 49, <ext-link xlink:href="https://doi.org/10.1029/2022GL100681" ext-link-type="DOI">10.1029/2022GL100681</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Kay et al.(2007)Kay, Jones, Crooks, Kjeldsen, and Fung</label><mixed-citation>Kay, A. L., Jones, D. A., Crooks, S. M., Kjeldsen, T. R., and Fung, C. F.: An investigation of site-similarity approaches to generalisation of a rainfall–runoff model, Hydrol. Earth Syst. Sci., 11, 500–515, <ext-link xlink:href="https://doi.org/10.5194/hess-11-500-2007" ext-link-type="DOI">10.5194/hess-11-500-2007</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Lei et al.(2025)Lei, Lin, Zheng, Fei, Yin, and Ren</label><mixed-citation>Lei, X., Lin, P., Zheng, H., Fei, W., Yin, Z., and Ren, H.: Systematic Analyses of the Meteorological Forcing and Process Parameterization Uncertainties in Modeling Runoff with Noah-MP for the Upper Brahmaputra River Basin, J. Hydrol., 653, 132686, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2025.132686" ext-link-type="DOI">10.1016/j.jhydrol.2025.132686</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Lei et al.(2026)</label><mixed-citation>Lei, X., Lin, H., Zheng, K., and Lin, P.: Process diagnostics of snowmelt runoff in global hydrological models: Part II – Are more complex models better?, EGUsphere [preprint], <ext-link xlink:href="https://doi.org/10.5194/egusphere-2025-6073" ext-link-type="DOI">10.5194/egusphere-2025-6073</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Li et al.(2022)Li, Bisht, and Leung</label><mixed-citation>Li, L., Bisht, G., and Leung, L. R.: Spatial heterogeneity effects on land surface modeling of water and energy partitioning, Geosci. Model Dev., 15, 5489–5510, <ext-link xlink:href="https://doi.org/10.5194/gmd-15-5489-2022" ext-link-type="DOI">10.5194/gmd-15-5489-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Liang et al.(1994)Liang, Lettenmaier, Wood, and Burges</label><mixed-citation>Liang, X., Lettenmaier, D. P., Wood, E. F., and Burges, S. J.: A Simple Hydrologically Based Model of Land Surface Water and Energy Fluxes for General Circulation Models, J. Geophys. Res.-Atmos., 99, 14415–14428, <ext-link xlink:href="https://doi.org/10.1029/94JD00483" ext-link-type="DOI">10.1029/94JD00483</ext-link>, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Lin et al.(2018)Lin, Yang, Gochis, Yu, Maidment, Somos-Valenzuela, and David</label><mixed-citation>Lin, P., Yang, Z.-L., Gochis, D. J., Yu, W., Maidment, D. R., Somos-Valenzuela, M. A., and David, C. H.: Implementation of a Vector-Based River Network Routing Scheme in the Community WRF-Hydro Modeling Framework for Flood Discharge Simulation, Environ. Model. Softw, 107, 1–11, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2018.05.018" ext-link-type="DOI">10.1016/j.envsoft.2018.05.018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Lin et al.(2019)Lin, Pan, Beck, Yang, Yamazaki, Frasson, David, Durand, Pavelsky, Allen, Gleason, and Wood</label><mixed-citation>Lin, P., Pan, M., Beck, H. E., Yang, Y., Yamazaki, D., Frasson, R., David, C. H., Durand, M., Pavelsky, T. M., Allen, G. H., Gleason, C. J., and Wood, E. F.: Global Reconstruction of Naturalized River Flows at 2.94 Million Reaches, Water Resour. Res., 55, 6499–6516, <ext-link xlink:href="https://doi.org/10.1029/2019WR025287" ext-link-type="DOI">10.1029/2019WR025287</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Lin et al.(2022)</label><mixed-citation>Lin, P., Pan, M., and Yang, Y.: Global Reconstruction of Naturalized River Discharge at 2.94 Million River Reaches (GRADES), National Tibetan Plateau/Third Pole Environment Data Center [data set], <ext-link xlink:href="https://doi.org/10.11888/Terre.tpdc.272898" ext-link-type="DOI">10.11888/Terre.tpdc.272898</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Martens et al.(2017)Martens, Miralles, Lievens, Van Der Schalie, De Jeu, Fernández-Prieto, Beck, Dorigo, and Verhoest</label><mixed-citation>Martens, B., Miralles, D. G., Lievens, H., van der Schalie, R., de Jeu, R. A. M., Fernández-Prieto, D., Beck, H. E., Dorigo, W. A., and Verhoest, N. E. C.: GLEAM v3: satellite-based land evaporation and root-zone soil moisture, Geosci. Model Dev., 10, 1903–1925, <ext-link xlink:href="https://doi.org/10.5194/gmd-10-1903-2017" ext-link-type="DOI">10.5194/gmd-10-1903-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Müller Schmied et al.(2024)Müller Schmied, Trautmann, Ackermann, Cáceres, Flörke, Gerdener, Kynast, Peiris, Schiebener, Schumacher, and Döll</label><mixed-citation>Müller Schmied, H., Trautmann, T., Ackermann, S., Cáceres, D., Flörke, M., Gerdener, H., Kynast, E., Peiris, T. A., Schiebener, L., Schumacher, M., and Döll, P.: The global water resources and use model WaterGAP v2.2e: description and evaluation of modifications and new features, Geosci. Model Dev., 17, 8817–8852, <ext-link xlink:href="https://doi.org/10.5194/gmd-17-8817-2024" ext-link-type="DOI">10.5194/gmd-17-8817-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Nash and Sutcliffe(1970)</label><mixed-citation>Nash, J. and Sutcliffe, J.: River Flow Forecasting through Conceptual Models Part I — A Discussion of Principles, J. Hydrol., 10, 282–290, <ext-link xlink:href="https://doi.org/10.1016/0022-1694(70)90255-6" ext-link-type="DOI">10.1016/0022-1694(70)90255-6</ext-link>, 1970.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Oleson et al.(2010)Oleson, Lawrence, Flanner, Kluzek, Levis, Swenson, Thornton, Dai, Decker, Dickinson, Feddema, Heald, Lamarque, Niu, Qian, Running, Sakaguchi, Slater, Stöckli, Wang, Yang, Zeng, and Zeng</label><mixed-citation>Oleson, K. W., Lawrence, D. M., Flanner, M. G., Kluzek, E., Levis, S., Swenson, S. C., Thornton, E., Dai, A., Decker, M., Dickinson, R., Feddema, J., Heald, C. L., Lamarque, J.-F., Niu, G.-Y., Qian, T., Running, S., Sakaguchi, K., Slater, A., Stöckli, R., Wang, A., Yang, L., Zeng, X., and Zeng, X.: Technical Description of Version 4.0 of the Community Land Model (CLM), <ext-link xlink:href="https://doi.org/10.5065/D6RR1W7M" ext-link-type="DOI">10.5065/D6RR1W7M</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Oudin et al.(2010)Oudin, Kay, Andréassian, and Perrin</label><mixed-citation>Oudin, L., Kay, A., Andréassian, V., and Perrin, C.: Are Seemingly Physically Similar Catchments Truly Hydrologically Similar?, Water Resour. Res., 46, 2009WR008887, <ext-link xlink:href="https://doi.org/10.1029/2009WR008887" ext-link-type="DOI">10.1029/2009WR008887</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Pokhrel et al.(2014)Pokhrel, Koirala, Kanae, and Oki</label><mixed-citation>Pokhrel, Y. N., Koirala, S., Kanae, S., and Oki, T.: Incorporation of Groundwater Pumping in a Global Land Surface Model with the Representation of Human Impacts, Water Resour. Res., <ext-link xlink:href="https://doi.org/10.1002/2014WR015602" ext-link-type="DOI">10.1002/2014WR015602</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Poulter et al.(2011)Poulter, Ciais, Hodson, Lischke, Maignan, Plummer, and Zimmermann</label><mixed-citation>Poulter, B., Ciais, P., Hodson, E., Lischke, H., Maignan, F., Plummer, S., and Zimmermann, N. E.: Plant functional type mapping for earth system models, Geosci. Model Dev., 4, 993–1010, <ext-link xlink:href="https://doi.org/10.5194/gmd-4-993-2011" ext-link-type="DOI">10.5194/gmd-4-993-2011</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Qin et al.(2020)Qin, Abatzoglou, Siebert, Huning, AghaKouchak, Mankin, Hong, Tong, Davis, and Mueller</label><mixed-citation>Qin, Y., Abatzoglou, J. T., Siebert, S., Huning, L. S., AghaKouchak, A., Mankin, J. S., Hong, C., Tong, D., Davis, S. J., and Mueller, N. D.: Agricultural Risks from Changing Snowmelt, Nat. Clim. Change, 10, 459–465, <ext-link xlink:href="https://doi.org/10.1038/s41558-020-0746-8" ext-link-type="DOI">10.1038/s41558-020-0746-8</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Schaphoff et al.(2018)Schaphoff, Von Bloh, Rammig, Thonicke, Biemans, Forkel, Gerten, Heinke, Jägermeyr, Knauer, Langerwisch, Lucht, Müller, Rolinski, and Waha</label><mixed-citation>Schaphoff, S., von Bloh, W., Rammig, A., Thonicke, K., Biemans, H., Forkel, M., Gerten, D., Heinke, J., Jägermeyr, J., Knauer, J., Langerwisch, F., Lucht, W., Müller, C., Rolinski, S., and Waha, K.: LPJmL4 – a dynamic global vegetation model with managed land – Part 1: Model description, Geosci. Model Dev., 11, 1343–1375, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-1343-2018" ext-link-type="DOI">10.5194/gmd-11-1343-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Schulz and De Jong(2004)</label><mixed-citation>Schulz, O. and de Jong, C.: Snowmelt and sublimation: field experiments and modelling in the High Atlas Mountains of Morocco, Hydrol. Earth Syst. Sci., 8, 1076–1089, <ext-link xlink:href="https://doi.org/10.5194/hess-8-1076-2004" ext-link-type="DOI">10.5194/hess-8-1076-2004</ext-link>, 2004.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Stacke and Hagemann(2021)</label><mixed-citation>Stacke, T. and Hagemann, S.: HydroPy (v1.0): a new global hydrology model written in Python, Geosci. Model Dev., 14, 7795–7816, <ext-link xlink:href="https://doi.org/10.5194/gmd-14-7795-2021" ext-link-type="DOI">10.5194/gmd-14-7795-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Strasser et al.(2008)Strasser, Bernhardt, Weber, Liston, and Mauser</label><mixed-citation>Strasser, U., Bernhardt, M., Weber, M., Liston, G. E., and Mauser, W.: Is snow sublimation important in the alpine water balance?, The Cryosphere, 2, 53–66, <ext-link xlink:href="https://doi.org/10.5194/tc-2-53-2008" ext-link-type="DOI">10.5194/tc-2-53-2008</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Sutanudjaja et al.(2018)Sutanudjaja, Van Beek, Wanders, Wada, Bosmans, Drost, Van Der Ent, De Graaf, Hoch, De Jong, Karssenberg, López López, Peßenteiner, Schmitz, Straatsma, Vannametee, Wisser, and Bierkens</label><mixed-citation>Sutanudjaja, E. H., van Beek, R., Wanders, N., Wada, Y., Bosmans, J. H. C., Drost, N., van der Ent, R. J., de Graaf, I. E. M., Hoch, J. M., de Jong, K., Karssenberg, D., López López, P., Peßenteiner, S., Schmitz, O., Straatsma, M. W., Vannametee, E., Wisser, D., and Bierkens, M. F. P.: PCR-GLOBWB 2: a 5 arcmin global hydrological and water resources model, Geosci. Model Dev., 11, 2429–2453, <ext-link xlink:href="https://doi.org/10.5194/gmd-11-2429-2018" ext-link-type="DOI">10.5194/gmd-11-2429-2018</ext-link>, 2018. </mixed-citation></ref>
      <ref id="bib1.bibx44"><label>Tang et al.(2023)Tang, Clark, Knoben, Liu, Gharari, Arnal, Beck, Wood, Newman, and Papalexiou</label><mixed-citation>Tang, G., Clark, M. P., Knoben, W. J. M., Liu, H., Gharari, S., Arnal, L., Beck, H. E., Wood, A. W., Newman, A. J., and Papalexiou, S. M.: The Impact of Meteorological Forcing Uncertainty on Hydrological Modeling: A Global Analysis of Cryosphere Basins, Water Resour. Res., 59, e2022WR033767, <ext-link xlink:href="https://doi.org/10.1029/2022WR033767" ext-link-type="DOI">10.1029/2022WR033767</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Tang et al.(2006)Tang, Oki, and Kanae</label><mixed-citation>Tang, Q., Oki, T., and Kanae, S.: A Distributed Biosphere Hydrological Model (Dbhm) for Large River Basin, Proc. Hydraul. Eng., 50, 37–42, <ext-link xlink:href="https://doi.org/10.2208/prohe.50.37" ext-link-type="DOI">10.2208/prohe.50.37</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>Torres-Rojas et al.(2022)Torres-Rojas, Vergopolan, Herman, and Chaney</label><mixed-citation>Torres-Rojas, L., Vergopolan, N., Herman, J. D., and Chaney, N. W.: Towards an Optimal Representation of Sub-grid Heterogeneity in Land Surface Models, Water Resour. Res., 58, e2022WR032233, <ext-link xlink:href="https://doi.org/10.1029/2022WR032233" ext-link-type="DOI">10.1029/2022WR032233</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Wieder et al.(2022)Wieder, Kennedy, Lehner, Musselman, Rodgers, Rosenbloom, Simpson, and Yamaguchi</label><mixed-citation>Wieder, W. R., Kennedy, D., Lehner, F., Musselman, K. N., Rodgers, K. B., Rosenbloom, N., Simpson, I. R., and Yamaguchi, R.: Pervasive Alterations to Snow-Dominated Ecosystem Functions under Climate Change, P. Natl. Acad. Sci. USA, 119, e2202393119, <ext-link xlink:href="https://doi.org/10.1073/pnas.2202393119" ext-link-type="DOI">10.1073/pnas.2202393119</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx48"><label>Yamazaki et al.(2019)Yamazaki, Ikeshima, Sosa, Bates, Allen, and Pavelsky</label><mixed-citation>Yamazaki, D., Ikeshima, D., Sosa, J., Bates, P. D., Allen, G. H., and Pavelsky, T. M.: MERIT Hydro: A High-Resolution Global Hydrography Map Based on Latest Topography Dataset, Water Resour. Res., 55, 5053–5073, <ext-link xlink:href="https://doi.org/10.1029/2019WR024873" ext-link-type="DOI">10.1029/2019WR024873</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx49"><label>Yin et al.(2023)</label><mixed-citation>Yin, Z., Lin, P., Riggs, R., Allen, G. H., Lei, X., Zheng, Z., and Cai, S.: A Synthesis of Global Streamflow characteristics, Hydrometeorology, and catchment Attributes (GSHA) for Large Sample River-Centric Studies V1.1, Version 1.3, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.10433905" ext-link-type="DOI">10.5281/zenodo.10433905</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx50"><label>Yin et al.(2024)Yin, Lin, Riggs, Allen, Lei, Zheng, and Cai</label><mixed-citation>Yin, Z., Lin, P., Riggs, R., Allen, G. H., Lei, X., Zheng, Z., and Cai, S.: A synthesis of Global Streamflow Characteristics, Hydrometeorology, and Catchment Attributes (GSHA) for large sample river-centric studies, Earth Syst. Sci. Data, 16, 1559–1587, <ext-link xlink:href="https://doi.org/10.5194/essd-16-1559-2024" ext-link-type="DOI">10.5194/essd-16-1559-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx51"><label>Yokohata et al.(2020)Yokohata, Kinoshita, Sakurai, Pokhrel, Ito, Okada, Satoh, Kato, Nitta, Fujimori, Felfelani, Masaki, Iizumi, Nishimori, Hanasaki, Takahashi, Yamagata, and Emori</label><mixed-citation>Yokohata, T., Kinoshita, T., Sakurai, G., Pokhrel, Y., Ito, A., Okada, M., Satoh, Y., Kato, E., Nitta, T., Fujimori, S., Felfelani, F., Masaki, Y., Iizumi, T., Nishimori, M., Hanasaki, N., Takahashi, K., Yamagata, Y., and Emori, S.: MIROC-INTEG-LAND version 1: a global biogeochemical land surface model with human water management, crop growth, and land-use change, Geosci. Model Dev., 13, 4713–4747, <ext-link xlink:href="https://doi.org/10.5194/gmd-13-4713-2020" ext-link-type="DOI">10.5194/gmd-13-4713-2020</ext-link>, 2020.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Process diagnostics of snowmelt runoff in global hydrological and land surface models – Part 1: A systematic evaluation across basins of increasing complexity</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Andreadis et al.(2013)Andreadis, Schumann, and
Pavelsky</label><mixed-citation>
      
Andreadis, K. M., Schumann, G. J.-P., and Pavelsky, T.: A Simple Global River
Bankfull Width and Depth Database: Data and Analysis Note, Water Resour. Res.,
49, 7164–7168, <a href="https://doi.org/10.1002/wrcr.20440" target="_blank">https://doi.org/10.1002/wrcr.20440</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Beck et al.(2017)Beck, van Dijk, de Roo, Dutra, Fink, Orth, and
Schellekens</label><mixed-citation>
      
Beck, H. E., van Dijk, A. I. J. M., de Roo, A., Dutra, E., Fink, G., Orth, R., and Schellekens, J.: Global evaluation of runoff from 10 state-of-the-art hydrological models, Hydrol. Earth Syst. Sci., 21, 2881–2903, <a href="https://doi.org/10.5194/hess-21-2881-2017" target="_blank">https://doi.org/10.5194/hess-21-2881-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Beck et al.(2019)Beck, Wood, Pan, Fisher, Miralles, van Dijk,
McVicar, and Adler</label><mixed-citation>
      
Beck, H. E., Wood, E. F., Pan, M., Fisher, C. K., Miralles, D. G., van Dijk, A.
I. J. M., McVicar, T. R., and Adler, R. F.: MSWEP V2 Global 3-Hourly
0.1° Precipitation: Methodology and Quantitative Assessment, B.
Am. Meteorol. Soc, 100, 473–500, <a href="https://doi.org/10.1175/BAMS-D-17-0138.1" target="_blank">https://doi.org/10.1175/BAMS-D-17-0138.1</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Best et al.(2011)Best, Pryor, Clark, Rooney, Essery, Ménard,
Edwards, Hendry, Porson, Gedney, Mercado, Sitch, Blyth, Boucher, Cox,
Grimmond, and Harding</label><mixed-citation>
      
Best, M. J., Pryor, M., Clark, D. B., Rooney, G. G., Essery, R. L. H., Ménard, C. B., Edwards, J. M., Hendry, M. A., Porson, A., Gedney, N., Mercado, L. M., Sitch, S., Blyth, E., Boucher, O., Cox, P. M., Grimmond, C. S. B., and Harding, R. J.: The Joint UK Land Environment Simulator (JULES), model description – Part 1: Energy and water fluxes, Geosci. Model Dev., 4, 677–699, <a href="https://doi.org/10.5194/gmd-4-677-2011" target="_blank">https://doi.org/10.5194/gmd-4-677-2011</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Burek et al.(2020)Burek, Satoh, Kahil, Tang, Greve, Smilovic,
Guillaumot, Zhao, and Wada</label><mixed-citation>
      
Burek, P., Satoh, Y., Kahil, T., Tang, T., Greve, P., Smilovic, M., Guillaumot, L., Zhao, F., and Wada, Y.: Development of the Community Water Model (CWatM v1.04) – a high-resolution hydrological model for global and regional assessment of integrated water resources management, Geosci. Model Dev., 13, 3267–3298, <a href="https://doi.org/10.5194/gmd-13-3267-2020" target="_blank">https://doi.org/10.5194/gmd-13-3267-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Chai et al.(2025)Chai, Miao, Gentine, Mudryk, Thackeray, Berghuijs,
Wu, Fan, Slater, Sun, and Zwiers</label><mixed-citation>
      
Chai, Y., Miao, C., Gentine, P., Mudryk, L., Thackeray, C. W., Berghuijs,
W. R., Wu, Y., Fan, X., Slater, L., Sun, Q., and Zwiers, F.: Constrained
Earth System Models Show a Stronger Reduction in Future Northern Hemisphere
Snowmelt Water, Nat. Clim. Change, 15, <a href="https://doi.org/10.1038/s41558-025-02308-y" target="_blank">https://doi.org/10.1038/s41558-025-02308-y</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Chen et al.(2021)Chen, Liu, Mao, Wang, Zeng, Chen, Wang, and
Chen</label><mixed-citation>
      
Chen, H., Liu, J., Mao, G., Wang, Z., Zeng, Z., Chen, A., Wang, K., and Chen,
D.: Intercomparison of Ten ISI-MIP Models in Simulating Discharges along
the Lancang-Mekong River Basin, Sci. Total Environ, 765, 144494,
<a href="https://doi.org/10.1016/j.scitotenv.2020.144494" target="_blank">https://doi.org/10.1016/j.scitotenv.2020.144494</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>David et al.(2011)David, Maidment, Niu, Yang, Habets, and
Eijkhout</label><mixed-citation>
      
David, C. H., Maidment, D. R., Niu, G.-Y., Yang, Z.-L., Habets, F., and
Eijkhout, V.: River Network Routing on the NHDPlus Dataset, J.
Hydrometeorol, 12, 913–934, <a href="https://doi.org/10.1175/2011JHM1345.1" target="_blank">https://doi.org/10.1175/2011JHM1345.1</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Dudley et al.(2017)Dudley, Hodgkins, McHale, Kolian, and
Renard</label><mixed-citation>
      
Dudley, R., Hodgkins, G., McHale, M., Kolian, M., and Renard, B.: Trends in
Snowmelt-Related Streamflow Timing in the Conterminous United States, J.
Hydrol., 547, 208–221, <a href="https://doi.org/10.1016/j.jhydrol.2017.01.051" target="_blank">https://doi.org/10.1016/j.jhydrol.2017.01.051</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Feng and Gleason(2024a)</label><mixed-citation>
      
Feng, D. and Gleason, C. J.: More Flow Upstream and Less Flow Downstream:
The Changing Form and Function of Global Rivers, Science, 386,
1305–1311, <a href="https://doi.org/10.1126/science.adl5728" target="_blank">https://doi.org/10.1126/science.adl5728</a>, 2024a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Feng and Gleason(2024b)</label><mixed-citation>
      
Feng, D. and Gleason, C.: Global River Discharge Reanalysis dataset (GRDR), Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.13951712" target="_blank">https://doi.org/10.5281/zenodo.13951712</a>, 2024b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Fenicia et al.(2014)Fenicia, Kavetski, Savenije, Clark, Schoups,
Pfister, and Freer</label><mixed-citation>
      
Fenicia, F., Kavetski, D., Savenije, H. H. G., Clark, M. P., Schoups, G.,
Pfister, L., and Freer, J.: Catchment Properties, Function, and Conceptual
Model Representation: Is There a Correspondence?, Hydrol. Processes, 28,
2451–2467, <a href="https://doi.org/10.1002/hyp.9726" target="_blank">https://doi.org/10.1002/hyp.9726</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Guimberteau et al.(2018)Guimberteau, Zhu, Maignan, Huang, Yue,
Dantec-Nédélec, Ottlé, Jornet-Puig, Bastos, Laurent, Goll,
Bowring, Chang, Guenet, Tifafi, Peng, Krinner, Ducharne, Wang, Wang, Wang,
Wang, Yin, Lauerwald, Joetzjer, Qiu, Kim, and
Ciais</label><mixed-citation>
      
Guimberteau, M., Zhu, D., Maignan, F., Huang, Y., Yue, C., Dantec-Nédélec, S., Ottlé, C., Jornet-Puig, A., Bastos, A., Laurent, P., Goll, D., Bowring, S., Chang, J., Guenet, B., Tifafi, M., Peng, S., Krinner, G., Ducharne, A., Wang, F., Wang, T., Wang, X., Wang, Y., Yin, Z., Lauerwald, R., Joetzjer, E., Qiu, C., Kim, H., and Ciais, P.: ORCHIDEE-MICT (v8.4.1), a land surface model for the high latitudes: model description and validation, Geosci. Model Dev., 11, 121–163, <a href="https://doi.org/10.5194/gmd-11-121-2018" target="_blank">https://doi.org/10.5194/gmd-11-121-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Guo et al.(2024)Guo, Hou, Yang, and Mcvicar</label><mixed-citation>
      
Guo, H., Hou, Y., Yang, Y., and Mcvicar, T. R.: Global Evaluation of
Simulated High and Low Flows from 23 Macroscale Models, J.
Hydrometeorol, 25, 425–443, <a href="https://doi.org/10.1175/JHM-D-23-0176.1" target="_blank">https://doi.org/10.1175/JHM-D-23-0176.1</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Gupta et al.(2009)Gupta, Kling, Yilmaz, and
Martinez</label><mixed-citation>
      
Gupta, H. V., Kling, H., Yilmaz, K. K., and Martinez, G. F.: Decomposition of
the Mean Squared Error and NSE Performance Criteria: Implications for
Improving Hydrological Modelling, J. Hydrol., 377, 80–91,
<a href="https://doi.org/10.1016/j.jhydrol.2009.08.003" target="_blank">https://doi.org/10.1016/j.jhydrol.2009.08.003</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Haddeland et al.(2011)Haddeland, Clark, Franssen, Ludwig, Voß,
Arnell, Bertrand, Best, Folwell, Gerten, Gomes, Gosling, Hagemann, Hanasaki,
Harding, Heinke, Kabat, Koirala, Oki, Polcher, Stacke, Viterbo, Weedon, and
Yeh</label><mixed-citation>
      
Haddeland, I., Clark, D. B., Franssen, W., Ludwig, F., Voß, F., Arnell,
N. W., Bertrand, N., Best, M., Folwell, S., Gerten, D., Gomes, S., Gosling,
S. N., Hagemann, S., Hanasaki, N., Harding, R., Heinke, J., Kabat, P.,
Koirala, S., Oki, T., Polcher, J., Stacke, T., Viterbo, P., Weedon, G. P.,
and Yeh, P.: Multimodel Estimate of the Global Terrestrial Water
Balance: Setup and First Results, J. Hydrometeorol, 12, 869–884,
<a href="https://doi.org/10.1175/2011JHM1324.1" target="_blank">https://doi.org/10.1175/2011JHM1324.1</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Han et al.(2024)Han, Liu, Woods, McVicar, Yang, Wang, Hou, Guo, Li,
and Yang</label><mixed-citation>
      
Han, J., Liu, Z., Woods, R., McVicar, T. R., Yang, D., Wang, T., Hou, Y., Guo,
Y., Li, C., and Yang, Y.: Streamflow Seasonality in a Snow-Dwindling World,
Nature, 629, 1075–1081, <a href="https://doi.org/10.1038/s41586-024-07299-y" target="_blank">https://doi.org/10.1038/s41586-024-07299-y</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Hanasaki et al.(2008)Hanasaki, Kanae, Oki, Masuda, Motoya, Shirakawa,
Shen, and Tanaka</label><mixed-citation>
      
Hanasaki, N., Kanae, S., Oki, T., Masuda, K., Motoya, K., Shirakawa, N., Shen, Y., and Tanaka, K.: An integrated model for the assessment of global water resources – Part 1: Model description and input meteorological forcing, Hydrol. Earth Syst. Sci., 12, 1007–1025, <a href="https://doi.org/10.5194/hess-12-1007-2008" target="_blank">https://doi.org/10.5194/hess-12-1007-2008</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Harper et al.(2023)Harper, Lamarche, Hartley, Peylin, Ottlé,
Bastrikov, San Martín, Bohnenstengel, Kirches, Boettcher, Shevchuk,
Brockmann, and Defourny</label><mixed-citation>
      
Harper, K. L., Lamarche, C., Hartley, A., Peylin, P., Ottlé, C., Bastrikov, V., San Martín, R., Bohnenstengel, S. I., Kirches, G., Boettcher, M., Shevchuk, R., Brockmann, C., and Defourny, P.: A 29-year time series of annual 300 m resolution plant-functional-type maps for climate models, Earth Syst. Sci. Data, 15, 1465–1499, <a href="https://doi.org/10.5194/essd-15-1465-2023" target="_blank">https://doi.org/10.5194/essd-15-1465-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Hersbach et al.(2020)Hersbach, Bell, Berrisford, Hirahara,
Horányi, Muñoz-Sabater, Nicolas, Peubey, Radu, Schepers, Simmons,
Soci, Abdalla, Abellan, Balsamo, Bechtold, Biavati, Bidlot, Bonavita,
De Chiara, Dahlgren, Dee, Diamantakis, Dragani, Flemming, Forbes, Fuentes,
Geer, Haimberger, Healy, Hogan, Hólm, Janisková, Keeley, Laloyaux,
Lopez, Lupu, Radnoti, de Rosnay, Rozum, Vamborg, Villaume, and
Thépaut</label><mixed-citation>
      
Hersbach, H., Bell, B., Berrisford, P., Hirahara, S., Horányi, A.,
Muñoz-Sabater, J., Nicolas, J., Peubey, C., Radu, R., Schepers, D.,
Simmons, A., Soci, C., Abdalla, S., Abellan, X., Balsamo, G., Bechtold, P.,
Biavati, G., Bidlot, J., Bonavita, M., De Chiara, G., Dahlgren, P., Dee, D.,
Diamantakis, M., Dragani, R., Flemming, J., Forbes, R., Fuentes, M., Geer,
A., Haimberger, L., Healy, S., Hogan, R. J., Hólm, E., Janisková, M.,
Keeley, S., Laloyaux, P., Lopez, P., Lupu, C., Radnoti, G., de Rosnay, P.,
Rozum, I., Vamborg, F., Villaume, S., and Thépaut, J.-N.: The ERA5
Global Reanalysis, Q. J. Roy. Meteor. Soc., 146, 1999–2049,
<a href="https://doi.org/10.1002/qj.3803" target="_blank">https://doi.org/10.1002/qj.3803</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Hou et al.(2023)Hou, Guo, Yang, and Liu</label><mixed-citation>
      
Hou, Y., Guo, H., Yang, Y., and Liu, W.: Global Evaluation of Runoff Simulation
from Climate, Hydrological and Land Surface Models, Water Resour. Res., 59,
e2021WR031817, <a href="https://doi.org/10.1029/2021WR031817" target="_blank">https://doi.org/10.1029/2021WR031817</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Hu et al.(2022)Hu, Chen, Chen, Zhou, Peng, Li, and
Sang</label><mixed-citation>
      
Hu, Z., Chen, D., Chen, X., Zhou, Q., Peng, Y., Li, J., and Sang, Y.:
CCHZ-DISO: A Timely New Assessment System for Data Quality or
Model Performance From Da Dao Zhi Jian, Geophys. Res. Lett, 49,
<a href="https://doi.org/10.1029/2022GL100681" target="_blank">https://doi.org/10.1029/2022GL100681</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Kay et al.(2007)Kay, Jones, Crooks, Kjeldsen, and
Fung</label><mixed-citation>
      
Kay, A. L., Jones, D. A., Crooks, S. M., Kjeldsen, T. R., and Fung, C. F.: An investigation of site-similarity approaches to generalisation of a rainfall–runoff model, Hydrol. Earth Syst. Sci., 11, 500–515, <a href="https://doi.org/10.5194/hess-11-500-2007" target="_blank">https://doi.org/10.5194/hess-11-500-2007</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Lei et al.(2025)Lei, Lin, Zheng, Fei, Yin, and
Ren</label><mixed-citation>
      
Lei, X., Lin, P., Zheng, H., Fei, W., Yin, Z., and Ren, H.: Systematic Analyses
of the Meteorological Forcing and Process Parameterization Uncertainties in
Modeling Runoff with Noah-MP for the Upper Brahmaputra River Basin,
J. Hydrol., 653, 132686, <a href="https://doi.org/10.1016/j.jhydrol.2025.132686" target="_blank">https://doi.org/10.1016/j.jhydrol.2025.132686</a>,
2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Lei et al.(2026)</label><mixed-citation>
      
Lei, X., Lin, H., Zheng, K., and Lin, P.: Process diagnostics of snowmelt runoff in global hydrological models: Part II – Are more complex models better?, EGUsphere [preprint], <a href="https://doi.org/10.5194/egusphere-2025-6073" target="_blank">https://doi.org/10.5194/egusphere-2025-6073</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Li et al.(2022)Li, Bisht, and Leung</label><mixed-citation>
      
Li, L., Bisht, G., and Leung, L. R.: Spatial heterogeneity effects on land surface modeling of water and energy partitioning, Geosci. Model Dev., 15, 5489–5510, <a href="https://doi.org/10.5194/gmd-15-5489-2022" target="_blank">https://doi.org/10.5194/gmd-15-5489-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Liang et al.(1994)Liang, Lettenmaier, Wood, and
Burges</label><mixed-citation>
      
Liang, X., Lettenmaier, D. P., Wood, E. F., and Burges, S. J.: A Simple
Hydrologically Based Model of Land Surface Water and Energy Fluxes for
General Circulation Models, J. Geophys. Res.-Atmos., 99, 14415–14428,
<a href="https://doi.org/10.1029/94JD00483" target="_blank">https://doi.org/10.1029/94JD00483</a>, 1994.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Lin et al.(2018)Lin, Yang, Gochis, Yu, Maidment, Somos-Valenzuela,
and David</label><mixed-citation>
      
Lin, P., Yang, Z.-L., Gochis, D. J., Yu, W., Maidment, D. R.,
Somos-Valenzuela, M. A., and David, C. H.: Implementation of a Vector-Based
River Network Routing Scheme in the Community WRF-Hydro Modeling
Framework for Flood Discharge Simulation, Environ. Model. Softw, 107, 1–11,
<a href="https://doi.org/10.1016/j.envsoft.2018.05.018" target="_blank">https://doi.org/10.1016/j.envsoft.2018.05.018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Lin et al.(2019)Lin, Pan, Beck, Yang, Yamazaki, Frasson, David,
Durand, Pavelsky, Allen, Gleason, and Wood</label><mixed-citation>
      
Lin, P., Pan, M., Beck, H. E., Yang, Y., Yamazaki, D., Frasson, R., David,
C. H., Durand, M., Pavelsky, T. M., Allen, G. H., Gleason, C. J., and Wood,
E. F.: Global Reconstruction of Naturalized River Flows at 2.94
Million Reaches, Water Resour. Res., 55, 6499–6516,
<a href="https://doi.org/10.1029/2019WR025287" target="_blank">https://doi.org/10.1029/2019WR025287</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Lin et al.(2022)</label><mixed-citation>
      
Lin, P., Pan, M., and Yang, Y.: Global Reconstruction of Naturalized River Discharge at 2.94 Million River Reaches (GRADES), National Tibetan Plateau/Third Pole Environment Data Center [data set], <a href="https://doi.org/10.11888/Terre.tpdc.272898" target="_blank">https://doi.org/10.11888/Terre.tpdc.272898</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Martens et al.(2017)Martens, Miralles, Lievens, Van Der Schalie,
De Jeu, Fernández-Prieto, Beck, Dorigo, and
Verhoest</label><mixed-citation>
      
Martens, B., Miralles, D. G., Lievens, H., van der Schalie, R., de Jeu, R. A. M., Fernández-Prieto, D., Beck, H. E., Dorigo, W. A., and Verhoest, N. E. C.: GLEAM v3: satellite-based land evaporation and root-zone soil moisture, Geosci. Model Dev., 10, 1903–1925, <a href="https://doi.org/10.5194/gmd-10-1903-2017" target="_blank">https://doi.org/10.5194/gmd-10-1903-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Müller Schmied et al.(2024)Müller Schmied, Trautmann,
Ackermann, Cáceres, Flörke, Gerdener, Kynast, Peiris, Schiebener,
Schumacher, and Döll</label><mixed-citation>
      
Müller Schmied, H., Trautmann, T., Ackermann, S., Cáceres, D., Flörke, M., Gerdener, H., Kynast, E., Peiris, T. A., Schiebener, L., Schumacher, M., and Döll, P.: The global water resources and use model WaterGAP v2.2e: description and evaluation of modifications and new features, Geosci. Model Dev., 17, 8817–8852, <a href="https://doi.org/10.5194/gmd-17-8817-2024" target="_blank">https://doi.org/10.5194/gmd-17-8817-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Nash and Sutcliffe(1970)</label><mixed-citation>
      
Nash, J. and Sutcliffe, J.: River Flow Forecasting through Conceptual Models
Part I — A Discussion of Principles, J. Hydrol., 10, 282–290,
<a href="https://doi.org/10.1016/0022-1694(70)90255-6" target="_blank">https://doi.org/10.1016/0022-1694(70)90255-6</a>, 1970.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Oleson et al.(2010)Oleson, Lawrence, Flanner, Kluzek, Levis, Swenson,
Thornton, Dai, Decker, Dickinson, Feddema, Heald, Lamarque, Niu, Qian,
Running, Sakaguchi, Slater, Stöckli, Wang, Yang, Zeng, and
Zeng</label><mixed-citation>
      
Oleson, K. W., Lawrence, D. M., Flanner, M. G., Kluzek, E., Levis, S., Swenson,
S. C., Thornton, E., Dai, A., Decker, M., Dickinson, R., Feddema, J., Heald,
C. L., Lamarque, J.-F., Niu, G.-Y., Qian, T., Running, S., Sakaguchi, K.,
Slater, A., Stöckli, R., Wang, A., Yang, L., Zeng, X., and Zeng, X.:
Technical Description of Version 4.0 of the Community Land Model
(CLM), <a href="https://doi.org/10.5065/D6RR1W7M" target="_blank">https://doi.org/10.5065/D6RR1W7M</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Oudin et al.(2010)Oudin, Kay, Andréassian, and
Perrin</label><mixed-citation>
      
Oudin, L., Kay, A., Andréassian, V., and Perrin, C.: Are Seemingly
Physically Similar Catchments Truly Hydrologically Similar?, Water Resour.
Res., 46, 2009WR008887, <a href="https://doi.org/10.1029/2009WR008887" target="_blank">https://doi.org/10.1029/2009WR008887</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Pokhrel et al.(2014)Pokhrel, Koirala, Kanae, and
Oki</label><mixed-citation>
      
Pokhrel, Y. N., Koirala, S., Kanae, S., and Oki, T.: Incorporation of
Groundwater Pumping in a Global Land Surface Model with the
Representation of Human Impacts, Water Resour. Res.,
<a href="https://doi.org/10.1002/2014WR015602" target="_blank">https://doi.org/10.1002/2014WR015602</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Poulter et al.(2011)Poulter, Ciais, Hodson, Lischke, Maignan,
Plummer, and Zimmermann</label><mixed-citation>
      
Poulter, B., Ciais, P., Hodson, E., Lischke, H., Maignan, F., Plummer, S., and Zimmermann, N. E.: Plant functional type mapping for earth system models, Geosci. Model Dev., 4, 993–1010, <a href="https://doi.org/10.5194/gmd-4-993-2011" target="_blank">https://doi.org/10.5194/gmd-4-993-2011</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Qin et al.(2020)Qin, Abatzoglou, Siebert, Huning, AghaKouchak,
Mankin, Hong, Tong, Davis, and Mueller</label><mixed-citation>
      
Qin, Y., Abatzoglou, J. T., Siebert, S., Huning, L. S., AghaKouchak, A.,
Mankin, J. S., Hong, C., Tong, D., Davis, S. J., and Mueller, N. D.:
Agricultural Risks from Changing Snowmelt, Nat. Clim. Change, 10, 459–465,
<a href="https://doi.org/10.1038/s41558-020-0746-8" target="_blank">https://doi.org/10.1038/s41558-020-0746-8</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Schaphoff et al.(2018)Schaphoff, Von Bloh, Rammig, Thonicke, Biemans,
Forkel, Gerten, Heinke, Jägermeyr, Knauer, Langerwisch, Lucht,
Müller, Rolinski, and Waha</label><mixed-citation>
      
Schaphoff, S., von Bloh, W., Rammig, A., Thonicke, K., Biemans, H., Forkel, M., Gerten, D., Heinke, J., Jägermeyr, J., Knauer, J., Langerwisch, F., Lucht, W., Müller, C., Rolinski, S., and Waha, K.: LPJmL4 – a dynamic global vegetation model with managed land – Part 1: Model description, Geosci. Model Dev., 11, 1343–1375, <a href="https://doi.org/10.5194/gmd-11-1343-2018" target="_blank">https://doi.org/10.5194/gmd-11-1343-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Schulz and De Jong(2004)</label><mixed-citation>
      
Schulz, O. and de Jong, C.: Snowmelt and sublimation: field experiments and modelling in the High Atlas Mountains of Morocco, Hydrol. Earth Syst. Sci., 8, 1076–1089, <a href="https://doi.org/10.5194/hess-8-1076-2004" target="_blank">https://doi.org/10.5194/hess-8-1076-2004</a>, 2004.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Stacke and Hagemann(2021)</label><mixed-citation>
      
Stacke, T. and Hagemann, S.: HydroPy (v1.0): a new global hydrology model written in Python, Geosci. Model Dev., 14, 7795–7816, <a href="https://doi.org/10.5194/gmd-14-7795-2021" target="_blank">https://doi.org/10.5194/gmd-14-7795-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Strasser et al.(2008)Strasser, Bernhardt, Weber, Liston, and
Mauser</label><mixed-citation>
      
Strasser, U., Bernhardt, M., Weber, M., Liston, G. E., and Mauser, W.: Is snow sublimation important in the alpine water balance?, The Cryosphere, 2, 53–66, <a href="https://doi.org/10.5194/tc-2-53-2008" target="_blank">https://doi.org/10.5194/tc-2-53-2008</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Sutanudjaja et al.(2018)Sutanudjaja, Van Beek, Wanders, Wada,
Bosmans, Drost, Van Der Ent, De Graaf, Hoch, De Jong, Karssenberg,
López López, Peßenteiner, Schmitz, Straatsma, Vannametee, Wisser,
and Bierkens</label><mixed-citation>
      
Sutanudjaja, E. H., van Beek, R., Wanders, N., Wada, Y., Bosmans, J. H. C., Drost, N., van der Ent, R. J., de Graaf, I. E. M., Hoch, J. M., de Jong, K., Karssenberg, D., López López, P., Peßenteiner, S., Schmitz, O., Straatsma, M. W., Vannametee, E., Wisser, D., and Bierkens, M. F. P.: PCR-GLOBWB 2: a 5&thinsp;arcmin global hydrological and water resources model, Geosci. Model Dev., 11, 2429–2453, <a href="https://doi.org/10.5194/gmd-11-2429-2018" target="_blank">https://doi.org/10.5194/gmd-11-2429-2018</a>, 2018.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>Tang et al.(2023)Tang, Clark, Knoben, Liu, Gharari, Arnal, Beck,
Wood, Newman, and Papalexiou</label><mixed-citation>
      
Tang, G., Clark, M. P., Knoben, W. J. M., Liu, H., Gharari, S., Arnal, L.,
Beck, H. E., Wood, A. W., Newman, A. J., and Papalexiou, S. M.: The
Impact of Meteorological Forcing Uncertainty on Hydrological
Modeling: A Global Analysis of Cryosphere Basins, Water Resour.
Res., 59, e2022WR033767, <a href="https://doi.org/10.1029/2022WR033767" target="_blank">https://doi.org/10.1029/2022WR033767</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Tang et al.(2006)Tang, Oki, and Kanae</label><mixed-citation>
      
Tang, Q., Oki, T., and Kanae, S.: A Distributed Biosphere Hydrological Model
(Dbhm) for Large River Basin, Proc. Hydraul. Eng., 50, 37–42,
<a href="https://doi.org/10.2208/prohe.50.37" target="_blank">https://doi.org/10.2208/prohe.50.37</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Torres-Rojas et al.(2022)Torres-Rojas, Vergopolan, Herman, and
Chaney</label><mixed-citation>
      
Torres-Rojas, L., Vergopolan, N., Herman, J. D., and Chaney, N. W.: Towards an
Optimal Representation of Sub-grid Heterogeneity in Land Surface Models,
Water Resour. Res., 58, e2022WR032233, <a href="https://doi.org/10.1029/2022WR032233" target="_blank">https://doi.org/10.1029/2022WR032233</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Wieder et al.(2022)Wieder, Kennedy, Lehner, Musselman, Rodgers,
Rosenbloom, Simpson, and Yamaguchi</label><mixed-citation>
      
Wieder, W. R., Kennedy, D., Lehner, F., Musselman, K. N., Rodgers, K. B.,
Rosenbloom, N., Simpson, I. R., and Yamaguchi, R.: Pervasive Alterations to
Snow-Dominated Ecosystem Functions under Climate Change, P. Natl. Acad.
Sci. USA, 119, e2202393119, <a href="https://doi.org/10.1073/pnas.2202393119" target="_blank">https://doi.org/10.1073/pnas.2202393119</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Yamazaki et al.(2019)Yamazaki, Ikeshima, Sosa, Bates, Allen, and
Pavelsky</label><mixed-citation>
      
Yamazaki, D., Ikeshima, D., Sosa, J., Bates, P. D., Allen, G. H., and Pavelsky,
T. M.: MERIT Hydro: A High-Resolution Global Hydrography Map Based on
Latest Topography Dataset, Water Resour. Res., 55, 5053–5073,
<a href="https://doi.org/10.1029/2019WR024873" target="_blank">https://doi.org/10.1029/2019WR024873</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Yin et al.(2023)</label><mixed-citation>
      
Yin, Z., Lin, P., Riggs, R., Allen, G. H., Lei, X., Zheng, Z., and Cai, S.: A Synthesis of Global Streamflow characteristics, Hydrometeorology, and catchment Attributes (GSHA) for Large Sample River-Centric Studies V1.1, Version 1.3, Zenodo [data set], <a href="https://doi.org/10.5281/zenodo.10433905" target="_blank">https://doi.org/10.5281/zenodo.10433905</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Yin et al.(2024)Yin, Lin, Riggs, Allen, Lei, Zheng, and
Cai</label><mixed-citation>
      
Yin, Z., Lin, P., Riggs, R., Allen, G. H., Lei, X., Zheng, Z., and Cai, S.: A synthesis of Global Streamflow Characteristics, Hydrometeorology, and Catchment Attributes (GSHA) for large sample river-centric studies, Earth Syst. Sci. Data, 16, 1559–1587, <a href="https://doi.org/10.5194/essd-16-1559-2024" target="_blank">https://doi.org/10.5194/essd-16-1559-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Yokohata et al.(2020)Yokohata, Kinoshita, Sakurai, Pokhrel, Ito,
Okada, Satoh, Kato, Nitta, Fujimori, Felfelani, Masaki, Iizumi, Nishimori,
Hanasaki, Takahashi, Yamagata, and Emori</label><mixed-citation>
      
Yokohata, T., Kinoshita, T., Sakurai, G., Pokhrel, Y., Ito, A., Okada, M., Satoh, Y., Kato, E., Nitta, T., Fujimori, S., Felfelani, F., Masaki, Y., Iizumi, T., Nishimori, M., Hanasaki, N., Takahashi, K., Yamagata, Y., and Emori, S.: MIROC-INTEG-LAND version 1: a global biogeochemical land surface model with human water management, crop growth, and land-use change, Geosci. Model Dev., 13, 4713–4747, <a href="https://doi.org/10.5194/gmd-13-4713-2020" target="_blank">https://doi.org/10.5194/gmd-13-4713-2020</a>, 2020.

    </mixed-citation></ref-html>--></article>
