<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "https://jats.nlm.nih.gov/nlm-dtd/publishing/3.0/journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0" article-type="research-article">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">HESS</journal-id><journal-title-group>
    <journal-title>Hydrology and Earth System Sciences</journal-title>
    <abbrev-journal-title abbrev-type="publisher">HESS</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Hydrol. Earth Syst. Sci.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7938</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/hess-30-4867-2026</article-id><title-group><article-title>Metrics that matter: objective functions and their impact on signature representation in conceptual hydrological models</article-title><alt-title>Metrics that matter</alt-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1 aff2">
          <name><surname>Wagener</surname><given-names>Peter</given-names></name>
          <email>peter.wagener@ucalgary.ca</email>
        <ext-link>https://orcid.org/0009-0005-5560-3698</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Knoben</surname><given-names>Wouter J. M.</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-8301-3787</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1">
          <name><surname>Schütze</surname><given-names>Niels</given-names></name>
          
        <ext-link>https://orcid.org/0000-0002-2376-528X</ext-link></contrib>
        <contrib contrib-type="author" corresp="no" rid="aff1 aff2">
          <name><surname>Spieler</surname><given-names>Diana</given-names></name>
          
        <ext-link>https://orcid.org/0000-0003-3713-9148</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Institute of Hydrology and Meteorology, TUD Dresden University of Technology, Dresden, Germany</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>Schulich School of Engineering, Department of Civil Engineering, University of Calgary, Calgary, Canada</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Peter Wagener (peter.wagener@ucalgary.ca)</corresp></author-notes><pub-date><day>4</day><month>August</month><year>2026</year></pub-date>
      
      <volume>30</volume>
      <issue>15</issue>
      <fpage>4867</fpage><lpage>4888</lpage>
      <history>
        <date date-type="received"><day>3</day><month>November</month><year>2025</year></date>
           <date date-type="rev-request"><day>27</day><month>November</month><year>2025</year></date>
           <date date-type="rev-recd"><day>8</day><month>June</month><year>2026</year></date>
           <date date-type="accepted"><day>7</day><month>July</month><year>2026</year></date>
      </history>
      <permissions>
        <copyright-statement>Copyright: © 2026 Peter Wagener et al.</copyright-statement>
        <copyright-year>2026</copyright-year>
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026.html">This article is available from https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026.html</self-uri><self-uri xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026.pdf">The full text article is available as a PDF file from https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026.pdf</self-uri>
      <abstract><title>Abstract</title>

      <p id="d2e113">Although objective functions (OFs) are widely discussed in the literature, many modelling studies still default to a few common metrics, without much consideration of their relative strengths and weaknesses. This paper systematically investigates the impact of OF choice on the representation of various streamflow characteristics across 47 conceptual models and 10 hydro-climatically diverse catchments selected from the CARAVAN dataset. We use eight different OFs for calibration, including the Kling–Gupta efficiency (KGE), Nash–Sutcliffe efficiency (NSE), a variant of each that uses a streamflow transformation, and four more recently proposed metrics. We further account for parameter uncertainty by retaining the top 500 parameter sets for each combination of model, basin and OF. We evaluate the representation of 15 hydrological signatures that capture a relevant selection of streamflow characteristics to determine generalisable strengths and weaknesses of individual OFs across different models and hydroclimatic conditions. A random forest importance analysis showed that OF choice is often more influential than model structure choice in determining signature representation, with catchment characteristics remaining the dominant control. While certain signatures, particularly those related to flow variability, are relatively insensitive to OF choice, other signatures exhibit large performance shifts across different OFs. As other studies have already noted, no single OF performs best across all signatures. Our consistent multi-model and multi-catchment experiment now provides generalisable guidance on which OF best reproduces which aspect of streamflow behaviour. This helps to better select objective functions most appropriate for a specific modelling purpose and provides insights into which metric combinations can cover a broader range of flow characteristics.</p>
  </abstract>
    
<funding-group>
<award-group id="gs1">
<funding-source>National Oceanic and Atmospheric Administration</funding-source>
<award-id>NA22NWS4320003</award-id>
</award-group>
</funding-group>
</article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <label>1</label><title>Introduction</title>
      <p id="d2e125">Setting up and running a hydrological model requires modellers to make many decisions. These may include the choice of a suitable model (structure), sensible parameter boundaries, an appropriate calibration period, and which method to use for forcing data interpolation. Selecting which performance metric(s) to employ as the objective function during model calibration is another critical decision where research has shown that different performance metrics can lead to substantially different model outcomes. <xref ref-type="bibr" rid="bib1.bibx43" id="text.1"/>, for example, tested six alternative calibration metrics and found marked differences in the resulting hydrographs. <xref ref-type="bibr" rid="bib1.bibx21" id="text.2"/> compared variants of the KGE metric for low-flow simulation, observing large differences in annual runoff estimates. Studies such as <xref ref-type="bibr" rid="bib1.bibx51" id="text.3"/> and <xref ref-type="bibr" rid="bib1.bibx63" id="text.4"/> further demonstrated that the choice of performance metric for calibration influences how well the model reproduces key hydrological signatures (i.e., streamflow characteristics). Similarly, <xref ref-type="bibr" rid="bib1.bibx47" id="text.5"/> and <xref ref-type="bibr" rid="bib1.bibx57" id="text.6"/> show the impact calibration choices can have on the portrayal of climate change impacts.</p>
      <p id="d2e147">The choice of calibration metric, therefore, plays a crucial role among the various decisions involved in model calibration, evaluation, and diagnostic analysis. It directly determines how model parameters are optimized and, consequently, how well the resulting simulations reproduce both observed data and underlying hydrological processes. Despite the well-established influence of objective function selection, and an extensive body of research discussing the individual strengths and weaknesses of various performance metrics <xref ref-type="bibr" rid="bib1.bibx58 bib1.bibx3 bib1.bibx11 bib1.bibx39 bib1.bibx48 bib1.bibx7 bib1.bibx24 bib1.bibx42" id="paren.7"><named-content content-type="pre">e.g.</named-content></xref>, <xref ref-type="bibr" rid="bib1.bibx29" id="text.8"/> conclude that: “in most hydrologic modelling studies error metrics are chosen based on familiarity, [and] without consideration of the relative strengths and weaknesses” in their review of more than 60 different performance metrics. And indeed, a quick and informal review of 60 modelling studies in the existing literature library of the authors supports this claim by showing that most studies provide little to no reasoning for their choice of objective function (Fig. <xref ref-type="fig" rid="F1"/>; details on review in the Supplement). The papers that gave reasoning mainly referred to general characteristics and perceptions of their metric of choice. Reasons we documented (see Supplement) can roughly be classified into various forms of “the KGE is a more balanced metric than the NSE”, “NSE is a common metric for high flows”, “logNSE/logKGE focus on low flows”, “NSE/KGE are widely used metrics” or “we use this metric because it is comparable with previous studies”.</p>

      <fig id="F1"><label>Figure 1</label><caption><p id="d2e162">Number of analysed studies published during different five-year periods and the percentage of these studies that give no – general – or elaborate reasoning on their choice of objective function. This data stems from an informal/non-exhaustive literature review of 60 studies taken from the authors' literature database. Details regarding the review are given in Supplement.</p></caption>
        <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f01.png"/>

      </fig>

      <p id="d2e172">We believe that part of the prevailing tendency to define objective functions in an often ad-hoc manner is the broad but scattered and often site- and/or model-specific research on the impacts of different performance metrics used as objective functions. Most good modelling practice guidelines <xref ref-type="bibr" rid="bib1.bibx61 bib1.bibx31 bib1.bibx32" id="paren.9"><named-content content-type="pre">e.g.</named-content></xref> simply advise to select an appropriate objective function that aligns with the purpose of the modelling study, but do not provide guidance on how to do so. The question of how to connect the choice for a suitable calibration metric to a specific modelling purpose thus remains unclear <xref ref-type="bibr" rid="bib1.bibx30" id="paren.10"/>. While several studies have investigated connections between objective function choice and how the calibrated models represent certain streamflow characteristics, they are often limited in the number of models they use <xref ref-type="bibr" rid="bib1.bibx26 bib1.bibx45 bib1.bibx51 bib1.bibx63" id="paren.11"><named-content content-type="pre">1 model in e.g.</named-content></xref> or signatures they consider <xref ref-type="bibr" rid="bib1.bibx5 bib1.bibx48 bib1.bibx57" id="paren.12"><named-content content-type="pre">less than 5 in e.g.</named-content></xref>. As Table <xref ref-type="table" rid="T1"/> highlights, studies with a similar focus to our study also mostly consider catchments of only one specific region or country and no study the authors are aware of tested more than 12 models while also considering several metrics and signatures.</p>

<table-wrap id="T1" specific-use="star"><label>Table 1</label><caption><p id="d2e198">Summary of similar studies that investigate the impact of objective function choice on the representation of hydrological signatures. We compare the number of investigated models, catchments, metrics and signatures to this study. Note that studies with an <sup>*</sup> asterisk also tested combinations of different metrics (i.e. composite criteria or multi-metric objective functions) in addition to individual calibration metrics.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="7">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="right"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Publication</oasis:entry>
         <oasis:entry colname="col2">Main Focus</oasis:entry>
         <oasis:entry colname="col3">Catchments</oasis:entry>
         <oasis:entry colname="col4">Location</oasis:entry>
         <oasis:entry colname="col5">Models</oasis:entry>
         <oasis:entry colname="col6">OFs</oasis:entry>
         <oasis:entry colname="col7">Signatures</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx3" id="text.13"/></oasis:entry>
         <oasis:entry colname="col2">OF impact on streamflow simulations</oasis:entry>
         <oasis:entry colname="col3">179</oasis:entry>
         <oasis:entry colname="col4">Brazil</oasis:entry>
         <oasis:entry colname="col5">1</oasis:entry>
         <oasis:entry colname="col6">11<sup>*</sup></oasis:entry>
         <oasis:entry colname="col7">7</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx5" id="text.14"/></oasis:entry>
         <oasis:entry colname="col2">OF impact on streamflow forecast</oasis:entry>
         <oasis:entry colname="col3">22</oasis:entry>
         <oasis:entry colname="col4">Chile</oasis:entry>
         <oasis:entry colname="col5">3</oasis:entry>
         <oasis:entry colname="col6">12<sup>*</sup></oasis:entry>
         <oasis:entry colname="col7">5</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx26" id="text.15"/></oasis:entry>
         <oasis:entry colname="col2">OF impact on signatures</oasis:entry>
         <oasis:entry colname="col3">33</oasis:entry>
         <oasis:entry colname="col4">Ireland</oasis:entry>
         <oasis:entry colname="col5">1</oasis:entry>
         <oasis:entry colname="col6">6<sup>*</sup></oasis:entry>
         <oasis:entry colname="col7">18</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx28" id="text.16"/></oasis:entry>
         <oasis:entry colname="col2">OF impact on signatures</oasis:entry>
         <oasis:entry colname="col3">1</oasis:entry>
         <oasis:entry colname="col4">ACF-Basin, USA</oasis:entry>
         <oasis:entry colname="col5">1</oasis:entry>
         <oasis:entry colname="col6">8<sup>*</sup></oasis:entry>
         <oasis:entry colname="col7">167</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx45" id="text.17"/></oasis:entry>
         <oasis:entry colname="col2">calibration impact on flood/drought</oasis:entry>
         <oasis:entry colname="col3">9</oasis:entry>
         <oasis:entry colname="col4">Thur Basin, Switzerland</oasis:entry>
         <oasis:entry colname="col5">1</oasis:entry>
         <oasis:entry colname="col6">2</oasis:entry>
         <oasis:entry colname="col7">6</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx47" id="text.18"/></oasis:entry>
         <oasis:entry colname="col2">calibration impact on climate projections</oasis:entry>
         <oasis:entry colname="col3">3</oasis:entry>
         <oasis:entry colname="col4">Colorado Basin, USA</oasis:entry>
         <oasis:entry colname="col5">4</oasis:entry>
         <oasis:entry colname="col6">4</oasis:entry>
         <oasis:entry colname="col7">6</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx48" id="text.19"/></oasis:entry>
         <oasis:entry colname="col2">OF impact on high flows</oasis:entry>
         <oasis:entry colname="col3">492</oasis:entry>
         <oasis:entry colname="col4">USA</oasis:entry>
         <oasis:entry colname="col5">2</oasis:entry>
         <oasis:entry colname="col6">5</oasis:entry>
         <oasis:entry colname="col7">2</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx51" id="text.20"/></oasis:entry>
         <oasis:entry colname="col2">OF impact on signatures</oasis:entry>
         <oasis:entry colname="col3">25</oasis:entry>
         <oasis:entry colname="col4">Tennessee, USA</oasis:entry>
         <oasis:entry colname="col5">1</oasis:entry>
         <oasis:entry colname="col6">29<sup>*</sup></oasis:entry>
         <oasis:entry colname="col7">13</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx57" id="text.21"/></oasis:entry>
         <oasis:entry colname="col2">OF impact on climate projections</oasis:entry>
         <oasis:entry colname="col3">37</oasis:entry>
         <oasis:entry colname="col4">Quebec, Canada</oasis:entry>
         <oasis:entry colname="col5">12</oasis:entry>
         <oasis:entry colname="col6">3</oasis:entry>
         <oasis:entry colname="col7">4</oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"><xref ref-type="bibr" rid="bib1.bibx63" id="text.22"/></oasis:entry>
         <oasis:entry colname="col2">OF impact on signatures</oasis:entry>
         <oasis:entry colname="col3">27</oasis:entry>
         <oasis:entry colname="col4">Tennessee, USA</oasis:entry>
         <oasis:entry colname="col5">1</oasis:entry>
         <oasis:entry colname="col6">9<sup>*</sup></oasis:entry>
         <oasis:entry colname="col7">12</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">This study</oasis:entry>
         <oasis:entry colname="col2">OF impact on signatures</oasis:entry>
         <oasis:entry colname="col3">10</oasis:entry>
         <oasis:entry colname="col4">worldwide</oasis:entry>
         <oasis:entry colname="col5">47</oasis:entry>
         <oasis:entry colname="col6">8</oasis:entry>
         <oasis:entry colname="col7">15</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e595">Some of the existing studies show recurring tendencies across regions and models. For instance, <xref ref-type="bibr" rid="bib1.bibx51" id="text.23"/> and <xref ref-type="bibr" rid="bib1.bibx26" id="text.24"/> found that traditional objective functions such as NSE or KGE can perform as well or better than tailored signature-based metrics for predicting flow characteristics not explicitly included in calibration. <xref ref-type="bibr" rid="bib1.bibx48" id="text.25"/> similarly noted that application-specific metrics may excel at their target but can degrade performance on other metrics. Generally, most studies can identify a preferred metric or metric combination. However, that selection is strongly conditioned to the modelling purpose, the flow conditions of interest, and the hydroclimatic setting of the individual study (e.g. seasonal forecasting in snow-influenced catchments for <xref ref-type="bibr" rid="bib1.bibx5" id="text.26"/> or tropical low-flow applications for <xref ref-type="bibr" rid="bib1.bibx3" id="text.27"/>). A robust transfer of findings to new contexts is therefore difficult and further complicated by inconsistent experimental designs across studies (differing catchment sets, model selections, and signature choices), which prevent a straightforward synthesis on how metric choices influence signature representation.</p>
      <p id="d2e613">To address this problem, this study aims to develop a more generalised understanding of how the choice of performance metric used for model calibration influences the representation of the hydrological behaviour. We focus on single-objective calibration because it remains the calibration standard and can provide insights into metric combinations that might serve as complementary multi-objective calibration targets. By systematically analysing catchments that span a broad range of climatic conditions and employing a large set of conceptual model structures, we seek to identify general benefits and limitations associated with different calibration metrics across diverse hydrological contexts. The selected set of catchments, though limited in number, was chosen based on differing climate indices (moisture, seasonality, and fraction of snow), to ensure a representative spread of hydroclimatic conditions. This enables us to examine how models calibrated with different performance metrics behave under varying hydrological regimes while also balancing depth with breadth of analysis <xref ref-type="bibr" rid="bib1.bibx25" id="paren.28"/>. We evaluate how 47 conceptual model structures, each calibrated using alternative performance metrics and an ensemble of high-performing parameter sets to account for parameter uncertainty, reproduce 15 selected hydrological signatures that describe distinct aspects of the flow regime, such as flow variability, baseflow contribution, and runoff dynamics. This multi-dimensional setup contributes to identifying generalisable strengths and weaknesses of individual performance metrics that extend beyond their known sensitivities to particular flow conditions, offering new insights into how calibration choices influence process realism in conceptual hydrological modelling. The following sections will introduce the catchments, models and signatures we used (Sect. <xref ref-type="sec" rid="Ch1.S2"/>), present the results (Sect. <xref ref-type="sec" rid="Ch1.S3"/>) and provide an overarching discussion of the implications metric choice can have on the representation of the hydrological behaviour as well as limitations and future steps (Sect. <xref ref-type="sec" rid="Ch1.S4"/>). Conclusions are presented in Sect. <xref ref-type="sec" rid="Ch1.S5"/>.</p>
</sec>
<sec id="Ch1.S2">
  <label>2</label><title>Methodology</title>
      <p id="d2e635">Figure <xref ref-type="fig" rid="F2"/> gives an overview of the methodology used in this study. We calibrate 47 conceptual hydrological models from the Modular Assessment of Rainfall–Runoff Models Toolbox (MARRMoT) <xref ref-type="bibr" rid="bib1.bibx38 bib1.bibx60" id="paren.29"/> using eight different calibration metrics. The models and metrics used are introduced in Sect. <xref ref-type="sec" rid="Ch1.S2.SS1.SSS1"/> and <xref ref-type="sec" rid="Ch1.S2.SS1.SSS2"/> respectively. We calibrate all of these models (as described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS2"/>) for a subset of 10 hydro-climatically differing catchments from the CARAVAN dataset <xref ref-type="bibr" rid="bib1.bibx41" id="paren.30"/> as introduced in Sect. <xref ref-type="sec" rid="Ch1.S2.SS1.SSS3"/>. For each calibration run, we retain the top 500 parameter sets for further investigation. After selecting only the well-performing model structures according to a benchmark procedure (Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>) we use the simulated discharge timeseries to calculate 15 selected signatures that represent varying aspects of the hydrological behaviour (as introduced in Sect. <xref ref-type="sec" rid="Ch1.S2.SS4"/>). Through a comparison of the simulated and observed signature values we are able to identify the strengths and weaknesses of the tested metrics in representing the hydrological regime over a broad range of hydro-climatic conditions and different model complexities.</p>

      <fig id="F2" specific-use="star"><label>Figure 2</label><caption><p id="d2e661">Overview of the experiment design of this study. <bold>(a)</bold> Shows the geographical location of the study catchments. <bold>(b)</bold> Visualizes the different complexities (number of storages vs number of parameters) of the investigated MARRMoT Models. <bold>(c)</bold> Lists the eight objective functions used for calibration. <bold>(d)</bold> Lists the 15 hydrological signatures investigated and shows a schematic visualization of some of them. </p></caption>
        <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f02.png"/>

      </fig>


<sec id="Ch1.S2.SS1">
  <label>2.1</label><title>Experiment Components</title>
<sec id="Ch1.S2.SS1.SSS1">
  <label>2.1.1</label><title>Models</title>
      <p id="d2e699">We obtained the 47 models used in this study from the MARRMoT toolbox <xref ref-type="bibr" rid="bib1.bibx38 bib1.bibx59" id="paren.31"/>, because the toolbox allows easy and consistent use of the different models. These models are based on the scientific literature and mimic some widely used and well-known models such as HBV, GR4J, TOPMODEL, VIC and HYMOD. A full list of all models included in this analysis can be found in the Supplement. Of note, only 12 of the 47 models possess a snow module. Details of these modules vary depending on the model under consideration. Figure <xref ref-type="fig" rid="F2"/>b shows how the model structures differ in their number of considered stores (i.e., state variables) and parameters and therefore gives an idea of the different models' complexity. The number of parameters ranges from one to 24 and the number of considered stores ranges from one to eight. The simplest model has one store and one parameter, whereas the most complex models have eight stores and 12 parameters, or three stores and 24 parameters. By using such a wide range of different complexities in model structures, we ensure that the strengths and weaknesses identified generalise across different model structures.</p>
</sec>
<sec id="Ch1.S2.SS1.SSS2">
  <label>2.1.2</label><title>Objective Functions</title>
      <p id="d2e715">We calibrate the 47 hydrological models with eight objective functions, each based on a different performance metric, and analyse how well the calibrated models simulate 15 streamflow signatures. We differentiate between model accuracy, expressed by the objective function value, and model adequacy, expressed by signature representation. This selection of eight metrics is necessarily non-exhaustive, many further metrics and transformations exist, but was chosen to span widely used and recently proposed formulations while keeping the analysis tractable. In general, performance metrics are numerical measures used to quantify model performance, whereas objective functions refer to the use of such metrics as optimization targets during model calibration. We therefore use “objective function” or “OF” when referring to their role in calibration, and “metric” when referring to their mathematical formulation or diagnostic properties. To describe them, we introduce the following set of common symbols: <inline-formula><mml:math id="M8" display="inline"><mml:mi>Q</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M9" display="inline"><mml:mi>P</mml:mi></mml:math></inline-formula> denote streamflow and precipitation respectively. <inline-formula><mml:math id="M10" display="inline"><mml:mi mathvariant="italic">σ</mml:mi></mml:math></inline-formula> and <inline-formula><mml:math id="M11" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula> are standard deviation and correlation. Subscripts are used to indicate the specifics of all mentioned variables, e.g. <inline-formula><mml:math id="M12" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M13" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> show simulated and observed streamflow respectively. Subscript <inline-formula><mml:math id="M14" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula> refers to individual time steps. For correlation, the subscripts <inline-formula><mml:math id="M15" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">pe</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M16" display="inline"><mml:mrow><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">sp</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> are used for Pearson and Spearman correlation. Overlines are used to indicate temporal means. In sums, the letter <inline-formula><mml:math id="M17" display="inline"><mml:mi>n</mml:mi></mml:math></inline-formula> is used to indicate the total length of the time series.</p>

<table-wrap id="T2" specific-use="star"><label>Table 2</label><caption><p id="d2e808">Equations of the eight performance metrics used as objective functions in this study. Numbers (1–8) correspond to metric references in the text.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="3">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">No.</oasis:entry>
         <oasis:entry colname="col2">Metric</oasis:entry>
         <oasis:entry colname="col3">Equation</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">(1)</oasis:entry>
         <oasis:entry colname="col2">Kling–Gupta Efficiency (KGE)</oasis:entry>
         <oasis:entry colname="col3">KGE <inline-formula><mml:math id="M18" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">pe</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">(2)</oasis:entry>
         <oasis:entry colname="col2">Nash–Sutcliffe Efficiency (NSE)</oasis:entry>
         <oasis:entry colname="col3">NSE <inline-formula><mml:math id="M19" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mo>(</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">(3)</oasis:entry>
         <oasis:entry colname="col2">Power-transformed KGE (KGE0.2)</oasis:entry>
         <oasis:entry colname="col3">KGE<inline-formula><mml:math id="M20" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mn mathvariant="normal">0.2</mml:mn></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mrow><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mover accent="true"><mml:mrow><mml:msubsup><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mn mathvariant="normal">0.2</mml:mn></mml:msubsup></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mover accent="true"><mml:mrow><mml:msubsup><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mrow><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:msubsup></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">0.2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">s</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mn mathvariant="normal">0.2</mml:mn><mml:mo>,</mml:mo><mml:mi mathvariant="normal">o</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi mathvariant="normal">pe</mml:mi><mml:mo>,</mml:mo><mml:mn mathvariant="normal">0.2</mml:mn></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(4)</oasis:entry>
         <oasis:entry colname="col2">logarithmic NSE (log NSE)</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M21" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mi>log⁡</mml:mi><mml:mo>(</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>+</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo>)</mml:mo><mml:mo>,</mml:mo><mml:mi mathvariant="italic">ϵ</mml:mi><mml:mo>&gt;</mml:mo><mml:mn mathvariant="normal">0</mml:mn></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1"/>
         <oasis:entry colname="col2"/>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M22" display="inline"><mml:mrow><mml:mi>log⁡</mml:mi><mml:mspace linebreak="nobreak" width="0.25em"/><mml:mi mathvariant="normal">NSE</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow><mml:mrow><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo stretchy="false" mathvariant="normal">̃</mml:mo></mml:mover><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo mathvariant="normal" stretchy="false">̃</mml:mo></mml:mover><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">(5)</oasis:entry>
         <oasis:entry colname="col2">Non-parametric KGE (KGE–NP)</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">KGE</mml:mi><mml:mi mathvariant="normal">NP</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>k</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mfenced close="|" open="|"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>I</mml:mi><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo>(</mml:mo><mml:mi>J</mml:mi><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mspace width="0.125em" linebreak="nobreak"/><mml:msub><mml:mover accent="true"><mml:mi>Q</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">sp</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">(6)</oasis:entry>
         <oasis:entry colname="col2">KGE-Split</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:mtext>KGE-Split</mml:mtext><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>Y</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>Y</mml:mi></mml:msubsup><mml:mfenced close=")" open="("><mml:mrow><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mrow><mml:mi mathvariant="normal">pe</mml:mi><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:mfenced></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">(7)</oasis:entry>
         <oasis:entry colname="col2">Signature-based Hydrologic Efficiency (SHE)</oasis:entry>
         <oasis:entry colname="col3">SHE <inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo><mml:msqrt><mml:mrow><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>/</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>/</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>P</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi>p</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">sp</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">(8)</oasis:entry>
         <oasis:entry colname="col2">Diagnostic Efficiency (DE)</oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:msub><mml:mi mathvariant="normal">DE</mml:mi><mml:mi mathvariant="normal">OF</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:msup><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mn mathvariant="normal">1</mml:mn><mml:mi>n</mml:mi></mml:mfrac></mml:mstyle><mml:msubsup><mml:mo>∑</mml:mo><mml:mrow><mml:mi>t</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow><mml:mi>n</mml:mi></mml:msubsup><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">s</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mrow><mml:mi mathvariant="normal">o</mml:mi><mml:mo>,</mml:mo><mml:mi>t</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:msup><mml:mfenced open="(" close=")"><mml:mrow><mml:msubsup><mml:mo>∫</mml:mo><mml:mn mathvariant="normal">0</mml:mn><mml:mn mathvariant="normal">1</mml:mn></mml:msubsup><mml:mfenced open="|" close="|"><mml:mrow><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msubsup><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi><mml:mo>↓</mml:mo></mml:msubsup><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo><mml:mo>-</mml:mo><mml:msubsup><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mo>↓</mml:mo></mml:msubsup><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow><mml:mrow><mml:msubsup><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi><mml:mo>↓</mml:mo></mml:msubsup><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow></mml:mfenced><mml:mspace linebreak="nobreak" width="0.125em"/><mml:mi mathvariant="normal">d</mml:mi><mml:mi>p</mml:mi></mml:mrow></mml:mfenced><mml:mn mathvariant="normal">2</mml:mn></mml:msup><mml:mo>+</mml:mo><mml:mo>(</mml:mo><mml:msub><mml:mi>r</mml:mi><mml:mi mathvariant="normal">pe</mml:mi></mml:msub><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:msup><mml:mo>)</mml:mo><mml:mn mathvariant="normal">2</mml:mn></mml:msup></mml:mrow></mml:msqrt></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e2024">The first two metrics are likely the most common objective functions used in hydrology <xref ref-type="bibr" rid="bib1.bibx46" id="paren.32"/>. That is, (1) the Kling-Gupta-Efficiency <xref ref-type="bibr" rid="bib1.bibx24" id="paren.33"><named-content content-type="pre">KGE;</named-content></xref> and (2) the Nash-Sutcliffe-Efficiency <xref ref-type="bibr" rid="bib1.bibx49" id="paren.34"><named-content content-type="pre">NSE;</named-content></xref> as defined in Table <xref ref-type="table" rid="T2"/>. Additionally, we use transformations of the first two metrics that aim to improve the representation of low flows. For KGE, we selected (3) a power transformation because of the known issues with logarithmic transformations <xref ref-type="bibr" rid="bib1.bibx58 bib1.bibx53" id="paren.35"/>. For NSE we selected (4) the logarithmic version because of its long-standing use <xref ref-type="bibr" rid="bib1.bibx42" id="paren.36"/>; here <inline-formula><mml:math id="M27" display="inline"><mml:mi mathvariant="italic">ϵ</mml:mi></mml:math></inline-formula> is a small positive offset added to streamflow to avoid undefined logarithms at zero flow. In addition, we include four recently developed metrics that intend to improve certain aspects of NSE or KGE or diagnose hydrological model behaviour.</p>
      <p id="d2e2057">First, this is (5) the non-parametric version of the KGE <xref ref-type="bibr" rid="bib1.bibx52" id="paren.37"><named-content content-type="pre">KGE-NP;</named-content></xref>. The major difference with the standard KGE is that the variability component is no longer based on the ratio of variances, but instead replaced by a flow-duration curve-based term. Additionally, the correlation term uses the Spearman rank correlation instead of the Pearson correlation. <xref ref-type="bibr" rid="bib1.bibx52" id="text.38"/> argue that the resulting benefit of their proposed metric is an overall better agreement between observations and simulations, except for high flows. Their reasoning is that more information is contained within the metric because the FDC is more complex than the standard deviation and the Spearman rank correlation leads to improvements for low flows. In the equation of (5) <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mi>I</mml:mi><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M29" display="inline"><mml:mrow><mml:mi>J</mml:mi><mml:mo>(</mml:mo><mml:mi>k</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> indicate time steps in which the <inline-formula><mml:math id="M30" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>th largest flow occurs in the simulated and observed time series respectively.</p>
      <p id="d2e2103">Second, another adaption of the KGE, (6) the KGE-Split <xref ref-type="bibr" rid="bib1.bibx18" id="paren.39"/> is evaluated. It is calculated like the regular KGE, but instead of calculating one value for the entire time series, a value for each year is calculated, and the mean over these values becomes the actual objective function value. <xref ref-type="bibr" rid="bib1.bibx18" id="text.40"/> developed it to put more emphasis on dry years, since the typically used “least squares” methods give more attention to high flows rather than low flows. The variables are identical to those used in the KGE, with <inline-formula><mml:math id="M31" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula> indicating the evaluated year and <inline-formula><mml:math id="M32" display="inline"><mml:mi>Y</mml:mi></mml:math></inline-formula> the total number of years as shown in Eq. (6).</p>
      <p id="d2e2126">Third, the (7) signature-based hydrologic efficiency <xref ref-type="bibr" rid="bib1.bibx33" id="paren.41"><named-content content-type="pre">SHE;</named-content></xref> normalizes both the bias and the variance term by the mean and variance of the precipitation. Like the KGE-NP, it also uses the Spearman description for the correlation term to improve on the low flow representations. The SHE as used in <xref ref-type="bibr" rid="bib1.bibx33" id="text.42"/> is defined in Eq. (7).</p>
      <p id="d2e2137">Fourth, we analyse (8) the Diagnostic Efficiency <xref ref-type="bibr" rid="bib1.bibx55" id="paren.43"><named-content content-type="pre">DE;</named-content></xref> where the bias term is replaced with a mean relative error and the variance term is the integral over all residuals of the relative error based on the flow duration curve. As the name suggests, this metric is designed as a diagnostic tool but will be applied here as an objective function. The general structure, with three parts for bias, variance, and timing, is similar to the KGE, but particularly the variance term proposes an interesting alternative to the KGE. To convert it into a similar objective function as the other candidates we use: DE<inline-formula><mml:math id="M33" display="inline"><mml:mrow><mml:msub><mml:mi/><mml:mi mathvariant="normal">OF</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>-</mml:mo></mml:mrow></mml:math></inline-formula> DE, the only metric for which we mark this calibration/analysis distinction in the notation since its two forms differ in formulation rather than in role alone. In Eq. (8), <inline-formula><mml:math id="M34" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula> denotes the exceedance probability and <inline-formula><mml:math id="M35" display="inline"><mml:mrow><mml:msup><mml:mi>Q</mml:mi><mml:mo>↓</mml:mo></mml:msup><mml:mo>(</mml:mo><mml:mi>p</mml:mi><mml:mo>)</mml:mo></mml:mrow></mml:math></inline-formula> the streamflow sorted along the flow-duration curve.</p>
      <p id="d2e2185">All of the eight used metrics can be disaggregated into three components representing bias, variability and correlation. For an easier overview of the differences between the individual metrics we summarized how they represent each component in Table <xref ref-type="table" rid="T3"/>.</p>

<table-wrap id="T3" specific-use="star"><label>Table 3</label><caption><p id="d2e2194">Overview of the three components of the eight metrics.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Metric</oasis:entry>
         <oasis:entry colname="col2">Bias</oasis:entry>
         <oasis:entry colname="col3">Variability</oasis:entry>
         <oasis:entry colname="col4">Correlation</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">KGE</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M36" display="inline"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mfrac></mml:mstyle></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M37" display="inline"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4">Pearson</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">NSE</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M38" display="inline"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mo>-</mml:mo><mml:mover accent="true"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow><mml:mo mathvariant="normal">‾</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3"><inline-formula><mml:math id="M39" display="inline"><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col4"><inline-formula><mml:math id="M40" display="inline"><mml:mrow><mml:mn mathvariant="normal">2</mml:mn><mml:mo>⋅</mml:mo><mml:mtext>Pearson</mml:mtext><mml:mo>⋅</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">s</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi mathvariant="italic">σ</mml:mi><mml:mi mathvariant="normal">o</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle></mml:mrow></mml:math></inline-formula></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">KGE0.2</oasis:entry>
         <oasis:entry colname="col2">as KGE</oasis:entry>
         <oasis:entry colname="col3">as KGE</oasis:entry>
         <oasis:entry colname="col4">as KGE</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">log NSE</oasis:entry>
         <oasis:entry colname="col2">as NSE</oasis:entry>
         <oasis:entry colname="col3">as NSE</oasis:entry>
         <oasis:entry colname="col4">as NSE</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">KGE-NP</oasis:entry>
         <oasis:entry colname="col2">as KGE</oasis:entry>
         <oasis:entry colname="col3">FDC-based</oasis:entry>
         <oasis:entry colname="col4">Spearman</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">KGE-Split</oasis:entry>
         <oasis:entry colname="col2">as KGE</oasis:entry>
         <oasis:entry colname="col3">as KGE</oasis:entry>
         <oasis:entry colname="col4">as KGE</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">SHE</oasis:entry>
         <oasis:entry colname="col2">as KGE</oasis:entry>
         <oasis:entry colname="col3">as KGE</oasis:entry>
         <oasis:entry colname="col4">Spearman</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">DE</oasis:entry>
         <oasis:entry colname="col2">Bias FDC-based</oasis:entry>
         <oasis:entry colname="col3">Residuals FDC-based</oasis:entry>
         <oasis:entry colname="col4">Pearson</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

</sec>
<sec id="Ch1.S2.SS1.SSS3">
  <label>2.1.3</label><title>Study Catchments</title>
      <p id="d2e2476">This study uses catchments from the CARAVAN database <xref ref-type="bibr" rid="bib1.bibx41" id="paren.44"/> which combines several large-sample datasets, include many CAMELS (Catchment Attributes and Meteorology for Large-sample Studies) versions of different countries or regions in a standardized way. While the latest update to CARAVAN now includes more than 20 000 different catchments <xref ref-type="bibr" rid="bib1.bibx20" id="paren.45"/> from many existing large-sample hydrology datasets, at the beginning of this study CARAVAN included 2901 catchments. At the time, CARAVAN included catchments in the United States <xref ref-type="bibr" rid="bib1.bibx1" id="paren.46"/>, Great Britain <xref ref-type="bibr" rid="bib1.bibx13" id="paren.47"/>, Brazil <xref ref-type="bibr" rid="bib1.bibx10" id="paren.48"/>, Chile <xref ref-type="bibr" rid="bib1.bibx4" id="paren.49"/> and Australia <xref ref-type="bibr" rid="bib1.bibx19" id="paren.50"/> from their respective CAMELS datasets, as well as the LamaH-CE <xref ref-type="bibr" rid="bib1.bibx35" id="paren.51"><named-content content-type="pre">Central Europe,</named-content></xref> and HYSETS <xref ref-type="bibr" rid="bib1.bibx6" id="paren.52"><named-content content-type="pre">North America,</named-content></xref> datasets. CARAVAN provides standardized meteorological forcing, streamflow data, and static catchment attributes (e.g., geophysical, sociological, climatological) with a median data length of 31 years. Our primary goal is to investigate the sensitivity of a large number of different model structures to objective function choice. To keep the analysis manageable (i.e., to “balance depth with breadth” <xref ref-type="bibr" rid="bib1.bibx25" id="paren.53"/>, we defined a subset of ten hydro-climatically differing catchments for our analysis. To determine these catchments, we used the climate classification procedure outlined in <xref ref-type="bibr" rid="bib1.bibx37" id="text.54"/> to quantify the aridity, seasonality and fraction snow in each of the CARAVAN basins. We then used a <inline-formula><mml:math id="M41" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-means clustering algorithm to divide the catchments into 10 clusters, and selected the catchment closest to each cluster centroid for use in this work. The location of the selected catchments is shown in Fig. 2a and a brief description of some catchment properties is given in Table <xref ref-type="table" rid="T4"/>. We also applied a threshold for maximum catchment area of 1000 km<sup>2</sup> to ensure a more direct influence between discharge, signatures and metrics, because for signatures like the Flashiness Index there might be a dampening effect in larger catchments. The full clustering procedure is described in more detail in  Sect. S2.2.</p>

<table-wrap id="T4" specific-use="star"><label>Table 4</label><caption><p id="d2e2539">Overview of catchment properties for the selected 10 CARAVAN basins. Details in Sect. S2.2.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="11">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="right"/>
     <oasis:colspec colnum="5" colname="col5" align="right"/>
     <oasis:colspec colnum="6" colname="col6" align="right"/>
     <oasis:colspec colnum="7" colname="col7" align="right"/>
     <oasis:colspec colnum="8" colname="col8" align="right"/>
     <oasis:colspec colnum="9" colname="col9" align="right"/>
     <oasis:colspec colnum="10" colname="col10" align="right"/>
     <oasis:colspec colnum="11" colname="col11" align="left"/>
     <oasis:thead>
       <oasis:row>
         <oasis:entry colname="col1">Cluster</oasis:entry>
         <oasis:entry colname="col2">Abbrev.</oasis:entry>
         <oasis:entry colname="col3">Dataset</oasis:entry>
         <oasis:entry colname="col4">Area</oasis:entry>
         <oasis:entry colname="col5">Moisture</oasis:entry>
         <oasis:entry colname="col6">Seasonality</oasis:entry>
         <oasis:entry colname="col7">Snow Frac</oasis:entry>
         <oasis:entry colname="col8">Precip</oasis:entry>
         <oasis:entry colname="col9">Runoff</oasis:entry>
         <oasis:entry colname="col10">Temp</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">–</oasis:entry>
         <oasis:entry colname="col2">–</oasis:entry>
         <oasis:entry colname="col3">–</oasis:entry>
         <oasis:entry colname="col4">(km<sup>2</sup>)</oasis:entry>
         <oasis:entry colname="col5">(–)</oasis:entry>
         <oasis:entry colname="col6">(–)</oasis:entry>
         <oasis:entry colname="col7">(–)</oasis:entry>
         <oasis:entry colname="col8">(mm yr<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col9">(mm yr<sup>−1</sup>)</oasis:entry>
         <oasis:entry colname="col10">(°C)</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">1</oasis:entry>
         <oasis:entry colname="col2">AUS1</oasis:entry>
         <oasis:entry colname="col3">CAMELS-AUS</oasis:entry>
         <oasis:entry colname="col4">125.74</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M46" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.39</oasis:entry>
         <oasis:entry colname="col6">0.59</oasis:entry>
         <oasis:entry colname="col7">0.00</oasis:entry>
         <oasis:entry colname="col8">837</oasis:entry>
         <oasis:entry colname="col9">117</oasis:entry>
         <oasis:entry colname="col10">18.3</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">2</oasis:entry>
         <oasis:entry colname="col2">BR6</oasis:entry>
         <oasis:entry colname="col3">CAMELS-BR</oasis:entry>
         <oasis:entry colname="col4">194.00</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M47" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.28</oasis:entry>
         <oasis:entry colname="col6">1.30</oasis:entry>
         <oasis:entry colname="col7">0.00</oasis:entry>
         <oasis:entry colname="col8">1451</oasis:entry>
         <oasis:entry colname="col9">490</oasis:entry>
         <oasis:entry colname="col10">22.5</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">3</oasis:entry>
         <oasis:entry colname="col2">AUS6</oasis:entry>
         <oasis:entry colname="col3">CAMELS-AUS</oasis:entry>
         <oasis:entry colname="col4">124.94</oasis:entry>
         <oasis:entry colname="col5"><inline-formula><mml:math id="M48" display="inline"><mml:mo>-</mml:mo></mml:math></inline-formula>0.04</oasis:entry>
         <oasis:entry colname="col6">1.67</oasis:entry>
         <oasis:entry colname="col7">0.00</oasis:entry>
         <oasis:entry colname="col8">680</oasis:entry>
         <oasis:entry colname="col9">166</oasis:entry>
         <oasis:entry colname="col10">15.4</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">4</oasis:entry>
         <oasis:entry colname="col2">C12</oasis:entry>
         <oasis:entry colname="col3">CAMELS</oasis:entry>
         <oasis:entry colname="col4">148.69</oasis:entry>
         <oasis:entry colname="col5">0.02</oasis:entry>
         <oasis:entry colname="col6">1.66</oasis:entry>
         <oasis:entry colname="col7">0.49</oasis:entry>
         <oasis:entry colname="col8">774</oasis:entry>
         <oasis:entry colname="col9">302</oasis:entry>
         <oasis:entry colname="col10">2.5</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5</oasis:entry>
         <oasis:entry colname="col2">C02</oasis:entry>
         <oasis:entry colname="col3">CAMELS</oasis:entry>
         <oasis:entry colname="col4">285.39</oasis:entry>
         <oasis:entry colname="col5">0.03</oasis:entry>
         <oasis:entry colname="col6">1.07</oasis:entry>
         <oasis:entry colname="col7">0.07</oasis:entry>
         <oasis:entry colname="col8">1120</oasis:entry>
         <oasis:entry colname="col9">362</oasis:entry>
         <oasis:entry colname="col10">10.9</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">6</oasis:entry>
         <oasis:entry colname="col2">GB3</oasis:entry>
         <oasis:entry colname="col3">CAMELS-GB</oasis:entry>
         <oasis:entry colname="col4">136.53</oasis:entry>
         <oasis:entry colname="col5">0.16</oasis:entry>
         <oasis:entry colname="col6">1.46</oasis:entry>
         <oasis:entry colname="col7">0.00</oasis:entry>
         <oasis:entry colname="col8">793</oasis:entry>
         <oasis:entry colname="col9">186</oasis:entry>
         <oasis:entry colname="col10">9.7</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">7</oasis:entry>
         <oasis:entry colname="col2">C03</oasis:entry>
         <oasis:entry colname="col3">CAMELS</oasis:entry>
         <oasis:entry colname="col4">127.83</oasis:entry>
         <oasis:entry colname="col5">0.33</oasis:entry>
         <oasis:entry colname="col6">0.85</oasis:entry>
         <oasis:entry colname="col7">0.08</oasis:entry>
         <oasis:entry colname="col8">1412</oasis:entry>
         <oasis:entry colname="col9">657</oasis:entry>
         <oasis:entry colname="col10">10.6</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">8</oasis:entry>
         <oasis:entry colname="col2">HYS</oasis:entry>
         <oasis:entry colname="col3">HYSETS</oasis:entry>
         <oasis:entry colname="col4">328.44</oasis:entry>
         <oasis:entry colname="col5">0.34</oasis:entry>
         <oasis:entry colname="col6">1.28</oasis:entry>
         <oasis:entry colname="col7">0.38</oasis:entry>
         <oasis:entry colname="col8">1211</oasis:entry>
         <oasis:entry colname="col9">661</oasis:entry>
         <oasis:entry colname="col10">3.6</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">9</oasis:entry>
         <oasis:entry colname="col2">GB2</oasis:entry>
         <oasis:entry colname="col3">CAMELS-GB</oasis:entry>
         <oasis:entry colname="col4">283.60</oasis:entry>
         <oasis:entry colname="col5">0.42</oasis:entry>
         <oasis:entry colname="col6">1.25</oasis:entry>
         <oasis:entry colname="col7">0.00</oasis:entry>
         <oasis:entry colname="col8">1209</oasis:entry>
         <oasis:entry colname="col9">665</oasis:entry>
         <oasis:entry colname="col10">8.0</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">10</oasis:entry>
         <oasis:entry colname="col2">LAM</oasis:entry>
         <oasis:entry colname="col3">LamaH-CE</oasis:entry>
         <oasis:entry colname="col4">102.29</oasis:entry>
         <oasis:entry colname="col5">0.58</oasis:entry>
         <oasis:entry colname="col6">0.58</oasis:entry>
         <oasis:entry colname="col7">0.32</oasis:entry>
         <oasis:entry colname="col8">1777</oasis:entry>
         <oasis:entry colname="col9">1300</oasis:entry>
         <oasis:entry colname="col10">0.9</oasis:entry>
         <oasis:entry colname="col11"/>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e3061">The 10 selected catchments span a wide range of climatic and hydrological settings. They differ substantially in aridity and mean temperature, ranging from humid, cool basins such as LAM and HYS (moisture index <inline-formula><mml:math id="M49" display="inline"><mml:mrow><mml:mo>≈</mml:mo><mml:mn mathvariant="normal">0.3</mml:mn></mml:mrow></mml:math></inline-formula>–0.6, mean temperatures near 0–4 °C) to warm catchments like AUS1 and BR6 (negative moisture indices indicating high aridity, mean temperature <inline-formula><mml:math id="M50" display="inline"><mml:mo>&gt;</mml:mo></mml:math></inline-formula> 18 °C). Snow fraction varies substantially, from snow‐free catchments in Australia and Brazil to snow‐influenced basins such as C12 and LAM. Precipitation and runoff also contrast strongly, with wetter, high‐runoff regions (e.g. LAM, GB2) versus more arid, low‐runoff sites (AUS1, BR6), illustrating the climatic and geographic diversity captured by the dataset and clustering approach. To avoid known issues with the ERA5-based potential evapotranspiration (PET) estimation from CARAVAN <xref ref-type="bibr" rid="bib1.bibx12" id="paren.55"/>, we are using the provided alternative based on Penman-Monteith. We have compared the CARAVAN data for precipitation, potential evapotranspiration and temperature against the native forcing data provided by each catchment's source dataset (e.g. CAMELS, CAMELS-AUS, CAMELS-BR, CAMELS-GB, HYSETS, or LamaH-CE) and found no systematic bias, confirming the CARAVAN forcing was of sufficient quality for consistent use across all catchments (details in Figs. S3–S12 in the Supplement).</p>
</sec>
</sec>
<sec id="Ch1.S2.SS2">
  <label>2.2</label><title>Model Calibration Procedure</title>
      <p id="d2e3093">All models were run on daily time step and calibrated with the Covariance Matrix Adaptation Evolution Strategy (CMA-ES) algorithm by <xref ref-type="bibr" rid="bib1.bibx27" id="text.56"/> as implemented in the MARRMoT Toolbox. We used a 10-year time period from 2004 to 2014 for calibration and a 10-year time period from 1993 to 2003 for evaluation with a one-year warm-up period each. The evaluation period of the catchment AUS6 had to be shortened to eight years (1981–1988) because of data gaps; the calibration period covers 1990–1999. We conducted 18 800 (<inline-formula><mml:math id="M51" display="inline"><mml:mrow><mml:mn mathvariant="normal">5</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">47</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">8</mml:mn><mml:mo>×</mml:mo><mml:mn mathvariant="normal">10</mml:mn></mml:mrow></mml:math></inline-formula>) calibration runs to calibrate each of the 47 models on all eight metrics in each of the ten catchments. To consider parameter uncertainty we calibrated each combination on five different seeds and retained the best 500 parameter sets across all seeds.</p>
      <p id="d2e3119">For all analysis, we focus on the calibration period only, because it provides the closest connection between objective function choice and resulting model behaviour. If objective functions cannot produce accurate signatures during calibration, it is unlikely that they will systematically do so during evaluation. We provide additional plots showing model and signature performance during the evaluation period in the Supplement (Figs. S13/S14).</p>
</sec>
<sec id="Ch1.S2.SS3">
  <label>2.3</label><title>Model Benchmarking</title>
      <p id="d2e3130">We establish a baseline performance benchmark to remove poorly performing models from further analysis. Based on earlier arguments and examples <xref ref-type="bibr" rid="bib1.bibx56 bib1.bibx54 bib1.bibx40 bib1.bibx36" id="paren.57"/>, we use an ensemble of benchmark models that can be used to establish the minimum performance a model should exceed. The HydroBM package <xref ref-type="bibr" rid="bib1.bibx36" id="paren.58"/>, provides 22 benchmark models that can easily be used. They cover three broad categories: benchmarks derived from streamflow data, such as the daily mean flow timeseries, which accounts for seasonality; benchmarks derived from rainfall-runoff ratios; and very simple empirical models that approximate aggregated catchment behaviour. For each catchment, we compute the performance of the timeseries these benchmark models create by comparing it to the observed discharge data in the calibration period. By calculating all eight performance metrics we can identify which benchmark model yields the highest performance for each metric and catchment. This performance will be used as the benchmark to beat. Model runs that fail to outperform this highest benchmark are excluded from further analysis, ensuring that only sufficiently reliable models are carried forward to the analysis of signature performance. For catchments with a snow fraction larger than 30 % we additionally remove all models without a snow module from the analysis should they manage to beat the benchmark.</p>
</sec>
<sec id="Ch1.S2.SS4">
  <label>2.4</label><title>Signature Selection</title>
      <p id="d2e3148">We selected 15 streamflow signatures to investigate the extent to which objective function choice influences a model's ability to replicate streamflow signatures. Our goal with this selection is to cover different aspects of the flow regime, as well as a range of hydrological processes <xref ref-type="bibr" rid="bib1.bibx44" id="paren.59"/>. The selected signatures are shown in Table <xref ref-type="table" rid="T5"/> and were all calculated using the TOSSH toolbox <xref ref-type="bibr" rid="bib1.bibx22" id="paren.60"/>, a more detailed description of each signature can be found in Sect. S2.5. Most of the considered signatures describe streamflow properties such as the slope of the flow duration curve (FDC Slope, 33rd–66th exceedance percentiles), the 5th and 95th streamflow percentiles (<inline-formula><mml:math id="M52" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M53" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>), the high and low flow duration (HFD &amp; LFD) and frequency (HFF &amp; LFF) as well as the mean half flow date (MHFD). The slope of the FDC describes the variability of the streamflow regime, indicating how quickly a river responds to rainfall events and how sustained the flows are during dry periods. Representing it correctly therefore indicates a good representation of in-catchment storage and flashiness behaviour. The 5th and 95th streamflow percentile describe the magnitude of low and high flows. Representing these signatures well indicates a good representation of extreme flow behaviour, while representing the high and low flow duration and frequency well ensures that also these important characteristics of extreme flow can be met. The MHFD helps to determine if the general timing and seasonality of streamflow is represented well. Additionally, we evaluate the impact of the objective function on the total runoff ratio (Total RR), event runoff ratio (Event RR), baseflow index (BFI), baseflow recession coefficient (BFRC), flashiness index (FI), variability index (VI) and rising limb density (RLD). These are intended to deliver information on the different catchment processes such as the general water balance (Total RR), individual storm responses (Event RR), the baseflow processes (BFI, BFRC) or general water storage behaviour and runoff dynamics (flashiness index, variability index, rising limb density).</p>

<table-wrap id="T5" specific-use="star"><label>Table 5</label><caption><p id="d2e3184">Overview of selected signatures.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="4">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:colspec colnum="3" colname="col3" align="left"/>
     <oasis:colspec colnum="4" colname="col4" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Signature</oasis:entry>
         <oasis:entry colname="col2">Abbreviation</oasis:entry>
         <oasis:entry colname="col3">Category</oasis:entry>
         <oasis:entry colname="col4">Units</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Slope of Flow Duration Curve (33–66)</oasis:entry>
         <oasis:entry colname="col2">FDC Slope</oasis:entry>
         <oasis:entry colname="col3">Flow Variability</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5th Streamflow Percentile</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M54" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">Low Flow Magnitude</oasis:entry>
         <oasis:entry colname="col4">mm d<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">95th Streamflow Percentile</oasis:entry>
         <oasis:entry colname="col2"><inline-formula><mml:math id="M56" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula></oasis:entry>
         <oasis:entry colname="col3">High Flow Magnitude</oasis:entry>
         <oasis:entry colname="col4">mm d<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Low Flow Frequency</oasis:entry>
         <oasis:entry colname="col2">LFF</oasis:entry>
         <oasis:entry colname="col3">Low Flow Occurrence</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">High Flow Frequency</oasis:entry>
         <oasis:entry colname="col2">HFF</oasis:entry>
         <oasis:entry colname="col3">High Flow Occurrence</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Low Flow Duration</oasis:entry>
         <oasis:entry colname="col2">LFD</oasis:entry>
         <oasis:entry colname="col3">Low Flow Duration</oasis:entry>
         <oasis:entry colname="col4">days per event</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">High Flow Duration</oasis:entry>
         <oasis:entry colname="col2">HFD</oasis:entry>
         <oasis:entry colname="col3">High Flow Duration</oasis:entry>
         <oasis:entry colname="col4">days per event</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Mean Half Flow Date</oasis:entry>
         <oasis:entry colname="col2">MHFD</oasis:entry>
         <oasis:entry colname="col3">Flow Timing</oasis:entry>
         <oasis:entry colname="col4">day of year</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Total Runoff Ratio</oasis:entry>
         <oasis:entry colname="col2">Total RR/TRR</oasis:entry>
         <oasis:entry colname="col3">Water Balance</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Event Runoff Ratio</oasis:entry>
         <oasis:entry colname="col2">Event RR/ERR</oasis:entry>
         <oasis:entry colname="col3">Partitioning/Connectivity</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Baseflow Index</oasis:entry>
         <oasis:entry colname="col2">BFI</oasis:entry>
         <oasis:entry colname="col3">Baseflow</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Baseflow Recession Coefficient</oasis:entry>
         <oasis:entry colname="col2">BFRC</oasis:entry>
         <oasis:entry colname="col3">Baseflow</oasis:entry>
         <oasis:entry colname="col4">1 d<sup>−1</sup></oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Flashiness Index</oasis:entry>
         <oasis:entry colname="col2">FI</oasis:entry>
         <oasis:entry colname="col3">Runoff Dynamics</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Variability Index</oasis:entry>
         <oasis:entry colname="col2">VI</oasis:entry>
         <oasis:entry colname="col3">Water Storage</oasis:entry>
         <oasis:entry colname="col4">–</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Rising Limb Density</oasis:entry>
         <oasis:entry colname="col2">RLD</oasis:entry>
         <oasis:entry colname="col3">Runoff Dynamics</oasis:entry>
         <oasis:entry colname="col4">rises d<sup>−1</sup></oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>


</sec>
<sec id="Ch1.S2.SS5">
  <label>2.5</label><title>Analysis of Objective Function Influence</title>
<sec id="Ch1.S2.SS5.SSS1">
  <label>2.5.1</label><title>Signature Error Metric</title>
      <p id="d2e3528">As most signature values have individual ranges and can differ substantially between catchments, we analyse normalized signature errors for better comparison. We first define the error at the level of individual retained runs, then aggregate it (additional details are provided in  Sect. S2.6).</p>
      <p id="d2e3531">For a single retained run (one catchment, one model, one objective function, one parameter set), the signature error is the deviation of the simulated signature from the observation,

              <disp-formula id="Ch1.E1" content-type="numbered"><label>9</label><mml:math id="M60" display="block"><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M61" display="inline"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the simulated signature value for model <inline-formula><mml:math id="M62" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>, catchment <inline-formula><mml:math id="M63" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>, objective function <inline-formula><mml:math id="M64" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>, and retained run <inline-formula><mml:math id="M65" display="inline"><mml:mi>r</mml:mi></mml:math></inline-formula>, and <inline-formula><mml:math id="M66" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the observed signature value for catchment <inline-formula><mml:math id="M67" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>.</p>
      <p id="d2e3640">To place signatures of different magnitudes on a comparable axis, errors are normalized on a per-catchment basis. The normalization denominator <inline-formula><mml:math id="M68" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is the largest absolute typical error across objective functions for that catchment,

              <disp-formula id="Ch1.E2" content-type="numbered"><label>10</label><mml:math id="M69" display="block"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>j</mml:mi></mml:msub><mml:mo>=</mml:mo><mml:munder><mml:mo movablelimits="false">max⁡</mml:mo><mml:mi>k</mml:mi></mml:munder><mml:mfenced open="(" close=")"><mml:mfenced close="|" open="|"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>e</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mfenced><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            where <inline-formula><mml:math id="M70" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>e</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></inline-formula> is the typical (median) error of the retained ensemble, obtained by first taking the median over the retained runs of each model and then the median across models,

              <disp-formula id="Ch1.E3" content-type="numbered"><label>11</label><mml:math id="M71" display="block"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>e</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mi mathvariant="normal">med</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>i</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mfenced close=")" open="("><mml:mrow><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mi mathvariant="normal">med</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>r</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

            The sequential median ensures that each model contributes equally to the typical error, independent of how many of its retained runs pass the benchmark.</p>
      <p id="d2e3760">The run-level normalized error is then

              <disp-formula id="Ch1.E4" content-type="numbered"><label>12</label><mml:math id="M72" display="block"><mml:mrow><mml:mi>e</mml:mi><mml:msub><mml:mi>m</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>e</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow><mml:mrow><mml:mstyle scriptlevel="+1"><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mo>max⁡</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi/><mml:mi>k</mml:mi></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mstyle><mml:mfenced close=")" open="("><mml:mfenced open="|" close="|"><mml:mrow><mml:mstyle scriptlevel="+1"><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mi mathvariant="normal">med</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi/><mml:mi>i</mml:mi></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mstyle><mml:mfenced close=")" open="("><mml:mrow><mml:mstyle scriptlevel="+1"><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mi mathvariant="normal">med</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msub><mml:mi/><mml:mi>r</mml:mi></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:mstyle><mml:mfenced open="(" close=")"><mml:mrow><mml:msub><mml:mi>S</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mi>y</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced></mml:mfenced></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>

            Positive values indicate overestimation and negative values underestimation of the observed signature. The normalization is performed catchment by catchment, so that an objective function that is the most biased in all catchments would reach <inline-formula><mml:math id="M73" display="inline"><mml:mrow><mml:mo>±</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> at the median level. Because <inline-formula><mml:math id="M74" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is set by the typical (median) error while individual runs may deviate further, run-level values can exceed <inline-formula><mml:math id="M75" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
      <p id="d2e3935">Aggregating over the retained ensemble gives a catchment- and objective-specific error metric. Since <inline-formula><mml:math id="M76" display="inline"><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> is independent of model and run, the median passes through the normalization,

              <disp-formula id="Ch1.E5" content-type="numbered"><label>13</label><mml:math id="M77" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="normal">EM</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mi mathvariant="normal">med</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>i</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mfenced open="(" close=")"><mml:mrow><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mi mathvariant="normal">med</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>r</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mfenced close=")" open="("><mml:mrow><mml:msub><mml:mi mathvariant="normal">em</mml:mi><mml:mrow><mml:mi>i</mml:mi><mml:mi>j</mml:mi><mml:mi>k</mml:mi><mml:mi>r</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:mfenced></mml:mrow></mml:mfenced><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>e</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mi>D</mml:mi><mml:mi>j</mml:mi></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>

            which by construction lies in <inline-formula><mml:math id="M78" display="inline"><mml:mrow><mml:mo>[</mml:mo><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>,</mml:mo><mml:mn mathvariant="normal">1</mml:mn><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>, with <inline-formula><mml:math id="M79" display="inline"><mml:mn mathvariant="normal">0</mml:mn></mml:math></inline-formula> indicating that the observed value is matched exactly.</p>
      <p id="d2e4048">Finally, aggregating across catchments yields a single value per signature and objective function, which is used to identify the objective function that best represents each signature,

              <disp-formula id="Ch1.E6" content-type="numbered"><label>14</label><mml:math id="M80" display="block"><mml:mrow><mml:msub><mml:mi mathvariant="normal">EM</mml:mi><mml:mrow><mml:mi>k</mml:mi><mml:mo>,</mml:mo><mml:mi mathvariant="normal">abs</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mi mathvariant="normal">med</mml:mi></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mi>j</mml:mi></mml:mtd></mml:mtr></mml:mtable><mml:mfenced close=")" open="("><mml:mrow><mml:mo fence="true">|</mml:mo><mml:msub><mml:mi mathvariant="normal">EM</mml:mi><mml:mrow><mml:mi>j</mml:mi><mml:mi>k</mml:mi></mml:mrow></mml:msub><mml:mo fence="true">|</mml:mo></mml:mrow></mml:mfenced><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
</sec>
<sec id="Ch1.S2.SS5.SSS2">
  <label>2.5.2</label><title>Statistical Testing of Objective Function Influence</title>
      <p id="d2e4101">To better understand which signatures are most affected by the objective function choice, we conducted a paired sample <inline-formula><mml:math id="M81" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-test: a pairwise comparison for two objective functions at a time by testing their median value of signature values across catchments. The null hypothesis for the <inline-formula><mml:math id="M82" display="inline"><mml:mi>t</mml:mi></mml:math></inline-formula>-test used here is <inline-formula><mml:math id="M83" display="inline"><mml:mrow><mml:msub><mml:mi>H</mml:mi><mml:mn mathvariant="normal">0</mml:mn></mml:msub><mml:mo>:</mml:mo></mml:mrow></mml:math></inline-formula> The mean values of the distributions are equal to each other. With 8 objective functions this leads to <inline-formula><mml:math id="M84" display="inline"><mml:mrow><mml:mo>(</mml:mo><mml:mstyle scriptlevel="+1"><mml:mtable class="substack"><mml:mtr><mml:mtd><mml:mn mathvariant="normal">8</mml:mn></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mn mathvariant="normal">2</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:mstyle><mml:mo>)</mml:mo><mml:mo>=</mml:mo><mml:mstyle displaystyle="false"><mml:mfrac style="text"><mml:mrow><mml:mn mathvariant="normal">8</mml:mn><mml:mo>⋅</mml:mo><mml:mn mathvariant="normal">7</mml:mn></mml:mrow><mml:mn mathvariant="normal">2</mml:mn></mml:mfrac></mml:mstyle><mml:mo>=</mml:mo><mml:mn mathvariant="normal">28</mml:mn></mml:mrow></mml:math></inline-formula> pairwise comparisons for each signature. The signature value entering each comparison is the per catchment median error EM<sub><italic>j</italic><italic>k</italic></sub> (Eq. <xref ref-type="disp-formula" rid="Ch1.E5"/>), in which model and parameter-set variability have already been collapsed through the sequential median (Sect. <xref ref-type="sec" rid="Ch1.S2.SS5.SSS1"/>); the test thus pairs the two objective functions catchment by catchment across the ten catchments, and the retained top-500 runs are not treated as independent samples.</p>
      <p id="d2e4185">This analysis highlights which signatures are most strongly influenced by the choice of objective function. We report the resulting <inline-formula><mml:math id="M86" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-values for each pairwise comparison and consider differences statistically significant at <inline-formula><mml:math id="M87" display="inline"><mml:mrow><mml:mi mathvariant="italic">α</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>. Not all distributions are expected to differ significantly, as some objective functions are likely to produce similar results. Based on these outcomes, we restrict subsequent analyses to signatures that are commonly significantly impacted by the choice of objective function, which we define as 40 % of the <inline-formula><mml:math id="M88" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-values being lower than 0.05.</p>
</sec>
<sec id="Ch1.S2.SS5.SSS3">
  <label>2.5.3</label><title>Random Forest for Attribute Importance</title>
      <p id="d2e4222">We apply random forest (RF) regression to assess the relative importance of catchment, objective function (OF), and model choice. We consider two complementary configurations. In the first, the response is the median simulated signature value across the retained ensemble per (model, catchment, OF) combination; in the second, the response is the corresponding median signature error against the observation (taken from Sect. <xref ref-type="sec" rid="Ch1.S2.SS5.SSS1"/>). In both cases the predictors are the categorical factor indices for model (<inline-formula><mml:math id="M89" display="inline"><mml:mi>i</mml:mi></mml:math></inline-formula>), OF (<inline-formula><mml:math id="M90" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>), and catchment (<inline-formula><mml:math id="M91" display="inline"><mml:mi>j</mml:mi></mml:math></inline-formula>).</p>
      <p id="d2e4248">We used the <monospace>TreeBagger</monospace> implementation in MATLAB with 500 trees and regression mode, extracting out-of-bag permuted predictor importance scores. Negative importance values, which indicate no predictive contribution, were set to zero. For each signature and configuration, importance values were normalized to sum to one, providing a direct measure of the relative influence of the three factors.</p>
</sec>
</sec>
</sec>
<sec id="Ch1.S3">
  <label>3</label><title>Results</title>
      <p id="d2e4264">The following sections will cover the outcomes of the benchmarking procedure (Sect. <xref ref-type="sec" rid="Ch1.S3.SS1"/>). Based on this, the general distribution of signature representation (Sect. <xref ref-type="sec" rid="Ch1.S3.SS2"/>) for (i) the detailed example of Total Runoff Ratio and (ii) all signatures (aggregated over the models) are shown. Afterwards, we will use the error metric (Eq. <xref ref-type="disp-formula" rid="Ch1.E5"/> from Sect. <xref ref-type="sec" rid="Ch1.S2.SS5.SSS1"/>) to highlight the skills and shortcomings of each objective function in Sect. <xref ref-type="sec" rid="Ch1.S3.SS3"/>. Further analysis is then used to test the statistical significance of objective function influence (Sect. <xref ref-type="sec" rid="Ch1.S3.SS4"/>) and assess the relative predictive importance of objective function choice (Sect. <xref ref-type="sec" rid="Ch1.S3.SS5"/>). Lastly, we use the aggregated error metric (Eq. <xref ref-type="disp-formula" rid="Ch1.E6"/> from Sect. <xref ref-type="sec" rid="Ch1.S2.SS5.SSS1"/>) to find the best performing objective functions regarding signature representation.</p>
<sec id="Ch1.S3.SS1">
  <label>3.1</label><title>Benchmarking the Models</title>
      <p id="d2e4293">Figure <xref ref-type="fig" rid="F3"/> shows the performance of all 47 calibrated models with their 500 best parameter sets (violins) compared to the highest calculated benchmark (red dash) as described in Sect. <xref ref-type="sec" rid="Ch1.S2.SS3"/>.</p>

      <fig id="F3" specific-use="star"><label>Figure 3</label><caption><p id="d2e4302">Benchmark Performance (driest catchment on left to wettest catchment on right). The distribution shows the objective function values across the top 500 parameter sets (across all calibrated seeds) for all 47 models and the red line shows the highest benchmark score of the objective function. The numbers below each violin indicate the number of models that outperform the benchmark with at least one of their parameter sets. The white circles indicate the median of the data in the violin.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f03.png"/>

        </fig>

      <p id="d2e4311">The number of models that beat the benchmark varies considerably across catchments and objective functions (number shown below each violin). In three catchments (C12, HYS, LAM) the number of models is comparatively low. The explanation is a strong snow influence, which produces seasonality that only the 12 of the 47 MARRMoT models that include a snow module can accurately reproduce. In other catchments, most models beat the benchmark, and the models that fail may be missing important process controls (e.g., groundwater representation) though this is more difficult to diagnose than the snow cases. In a handful of cases all models can beat the benchmark (e.g. GB2 for the KGE objective function), and for every combination of catchment and OF, there are at least 3 models exceeding the benchmark ensemble. As may be expected, more complex models (more parameters and stores), tend to outperform the benchmark more often and were retained more frequently (result not shown).</p>
</sec>
<sec id="Ch1.S3.SS2">
  <label>3.2</label><title>How accurate are the signature representations when different objective functions are used?</title>
      <p id="d2e4322">To begin the analysis of how objective functions affect hydrological adequacy, we compare the simulated signature values with the observed signature values for all eight objective functions. Figure <xref ref-type="fig" rid="F4"/> shows the distribution of absolute errors (<inline-formula><mml:math id="M92" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis) over all catchments (<inline-formula><mml:math id="M93" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) for each metric (subplots) for one of the 15 signatures, the Total Runoff Ratio. The plots for all other signatures can be found in Sect. S2.9 (Figs. S16 through S29).</p>

      <fig id="F4" specific-use="star"><label>Figure 4</label><caption><p id="d2e4343">Signature error (simulated – observed) on Total RR values for all objective functions. The shown distributions include all models and their 500 best parameter sets that passed the benchmark. The numbers above the violins indicate the number of runs (model <inline-formula><mml:math id="M94" display="inline"><mml:mo>×</mml:mo></mml:math></inline-formula> parameter set) shown within each violin. A number of 23 500 would indicate all models and parameter sets were retained.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f04.png"/>

        </fig>

      <p id="d2e4359">In Fig. <xref ref-type="fig" rid="F4"/>, there are three things of note. First, the location of the violin plots shows whether the signature values are underestimated (negative values), overestimated (positive values) or relatively unbiased (centered around zero). Second, the size of the violins indicates whether the signature predictions are well-constrained (short violin) or not (long violin). Third, the catchments are ordered from driest (left) to wettest (right) indicating if signature error correlates with climate.</p>
      <p id="d2e4365">There is no clear pattern in Runoff Ratio errors related to the catchments' aridity values. Instead, there appear to be two categories of results: objective functions that return mostly unbiased models in all catchments (KGE and KGE-NP), and objective functions where the signature representation of the resulting models varies strongly per catchment (e.g. NSE: relatively unbiased in some basins, large variability in signature errors in others). The KGE-NP constrains the variability within the models much better and is therefore assessed as the best available objective function for the general water balance as represented through the Total RR signature. For the remaining six objective functions, the Runoff Ratio tends to be underestimated and positive errors are relatively rare. This means that generally the models tend to put too much water into storage and/or evaporation or other sink terms, which may become relevant when modelling climate change impacts.</p>

      <fig id="F5" specific-use="star"><label>Figure 5</label><caption><p id="d2e4370">Comparison of signature representation capabilities across the different objective functions. The coloured points represent the median signature (across all retained models calibrated with a specific OF). The black dashes indicate the observed signature values. The grey violins represent the variability of the modelled signature values across the retained models with their top 500 parameter sets. In catchment AUS1, the observed value for the FDC slope is NaN, because the observed flow is 0 mm d<sup>−1</sup> for more than 66 % of the time steps in calibration period.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f05.png"/>

        </fig>

      <p id="d2e4391">To show the results for more signatures in an aggregated way, Fig. <xref ref-type="fig" rid="F5"/> shows the median simulated signature value over all models above the benchmark (underlying data can be seen in Figs. S16–S29). The median was chosen to achieve the most representable summary of the model ensemble, which is not susceptible to outliers. This gives an overall impression on the capability of the different models to represent the hydrological signatures of interest. It also allows a first impression of the impact different objective functions have on signature representation. In Fig. <xref ref-type="fig" rid="F5"/>, the coloured dots indicate the median modelled signature value for each objective function, while the black line shows the signature values calculated from observations. The grey violins indicate the variability in the modelled signature value if all individual model results with their 500 best parameter sets are considered (every model that beats the benchmark for any objective function). The plot therefore allows insights on how well a signature value can generally be represented through the different tested models and gives a first impression on the influence the objective function has on these representations. The <inline-formula><mml:math id="M96" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis denotes the 10 selected catchments and the <inline-formula><mml:math id="M97" display="inline"><mml:mi>y</mml:mi></mml:math></inline-formula>-axis gives the signature value in its native unit (see Table <xref ref-type="table" rid="T5"/>).</p>
      <p id="d2e4414">Initially, we can use this plot to compare the violin ranges to the observed values of signatures. For the vast majority of cases, the observed value is always within the boundaries of the model ensemble, which indicates that the tested conceptual models are able to capture the hydrologic variability across diverse catchments. The spread of the coloured dots around the observed values indicates different degrees of accuracy in the signature representation depending on the calibration metric used. For some signatures, the range of signature values is notably smaller (e.g. Total RR) than for others (e.g. Rising Limb Density). There are signatures (e.g. <inline-formula><mml:math id="M98" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, BFI) where the objective functions differ strongly from each other and signatures where different objective functions lead to similar modelled values (e.g. MHFD, HFD). To move beyond these very broad conclusions, the next section will use the previously introduced error metric (Eq. <xref ref-type="disp-formula" rid="Ch1.E5"/> from Sect. <xref ref-type="sec" rid="Ch1.S2.SS5.SSS1"/>) to assess the relative skill of each objective function in more detail.</p>
</sec>
<sec id="Ch1.S3.SS3">
  <label>3.3</label><title>Which objective functions show strengths or weaknesses for a specific hydrological signature?</title>
      <p id="d2e4440">In this section we use Eq. (<xref ref-type="disp-formula" rid="Ch1.E5"/>) to normalize the signature errors using the observed signature value. Figure <xref ref-type="fig" rid="F6"/> shows the normalized error in signature representation for all objective functions. Generally, we can see that the performance varies largely depending both on the signature (each violin in a plot) and catchments/models (within the violins).</p>

      <fig id="F6" specific-use="star"><label>Figure 6</label><caption><p id="d2e4449">Error Metric as given in Eq. (<xref ref-type="disp-formula" rid="Ch1.E5"/>) for all Objective Functions. The error metric is calculated for each objective function (subplot) per signature (<inline-formula><mml:math id="M99" display="inline"><mml:mi>x</mml:mi></mml:math></inline-formula>-axis) and catchment (data in violin). A value of 0 indicates no error in signature representation. A value of <inline-formula><mml:math id="M100" display="inline"><mml:mrow><mml:mo>-</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> or <inline-formula><mml:math id="M101" display="inline"><mml:mrow><mml:mo>+</mml:mo><mml:mn mathvariant="normal">1</mml:mn></mml:mrow></mml:math></inline-formula> would indicate a performance equal to the most extreme over/underestimation for the median across models and runs. As we are showing the top 500 runs for all benchmark-exceeding models explicitly, the values can exceed these boundaries. A number of 235 000 would indicate all models and parameter sets were retained.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f06.png"/>

        </fig>

      <p id="d2e4487">The KGE in Fig. <xref ref-type="fig" rid="F6"/> has good performance for the Total Runoff Ratio and the high flow percentile <inline-formula><mml:math id="M102" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> (violin has a smaller spread around the zero error line). This does not only apply to the median performance (white dot in the violin) but also has a high consistency among catchments and parameter sets (violin spread). However, particularly for the low flow percentile <inline-formula><mml:math id="M103" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, LFF, FDC slope, and Variability Index the representation obtained from calibrating to KGE relative to the other objective functions is often worse.</p>
      <p id="d2e4515">The NSE performs mediocre on almost all of the evaluated signatures: the range within catchments and runs (spread of each violin) is often large. This indicates that the suitability of the NSE metric can be very catchment- and model-dependent. HFF, FI, RLD and <inline-formula><mml:math id="M104" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> are usually underestimated by models calibrated on the NSE. Compared to the KGE, the performance on <inline-formula><mml:math id="M105" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, and FDC slope is typically better.</p>
      <p id="d2e4540">For KGE0.2 and log NSE, we see similar patterns in signature error. Both (KGE0.2 and log NSE) underestimate the runoff generation and high flows as well as the reactivity of the catchment (Flashiness Index). The log NSE is often the worst objective function among the eight investigated, indicated by high or low values. Both metrics show improvements compared to the KGE for FDC Slope and <inline-formula><mml:math id="M106" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, but not necessarily for low flow duration (LFD) and frequency (LFF). It is important to notice that the FDC Slope can also be met well if the runoff is systematically biased, which is the case with underestimation here. In many of the evaluated 15 signatures, the KGE0.2 shows a better performance than the log NSE (violin closer to 0).</p>
      <p id="d2e4554">The diagnostic efficiency generally performs similarly to the KGE0.2. It mainly improves representation of the BFI but has comparable issues for Total RR and <inline-formula><mml:math id="M107" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. Compared to the KGE, the errors are constrained better across catchments (smaller violins). In conclusion, the diagnostic efficiency proves to be a viable alternative for low-flow calibration, while offering benefits on additional signatures (such as BFI and VI).</p>
      <p id="d2e4568">The non-parametric version of the KGE performs worse than the KGE on high flows, but improves other signatures strongly, such as the BFI, FDC Slope, <inline-formula><mml:math id="M108" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and LFF. The KGE-NP is among the objective functions that perform best overall. The expectations (general improvement due to more information being included) for KGE-NP are thus largely met, with the main benefit being better incorporation of flow variability for a larger fraction of flow percentiles.</p>
      <p id="d2e4582">KGE-Split performs similar to the regular KGE, though shows a tendency towards larger errors. Like the NSE, it has no signature that is represented both more accurately and consistently than alternative objective functions. Compared to the regular KGE, there are slight improvements on selected signatures (FDC Slope, BFI, Variability Index).</p>
      <p id="d2e4585">Lastly, the SHE is one of the objective functions with the largest range in signature representation. The general patterns are similar to the KGE (as expected due to their similar formulation), except for selected signatures like FDC Slope. This indicates that a change from value-based correlation to rank-based can be influential for signature representation.</p>
      <p id="d2e4589">The violin plots also show that some signatures are affected very little by the choice of objective function such as the Rising Limb Density and the Mean Half Flow Date, which is indicated by a similar distribution among objective functions. This suggests that either the models or the catchment mainly influence the results. As the signature value cannot be simulated well from any of the models, this shows the limits of objective function influence.</p>
      <p id="d2e4592">The Supplement (Figs. S30 and S31) includes plots on the analysis with 25 and 100 parameter sets in comparison to the 500 used here. Since the results remain comparable, we assume that parameter uncertainty plays only a minor role here. Additionally, the SM also provides an example of the error metric collapsed to one value per catchment as it is being used in the assessment of overall skill (Fig. S32), which emphasizes the differences more, but does not account for model- and run-level uncertainty.</p>
</sec>
<sec id="Ch1.S3.SS4">
  <label>3.4</label><title>Does the objective function choice significantly affect signature representation?</title>
      <p id="d2e4604">This previous section showed the skills and shortcomings of the individual objective functions regarding signature representation. However, it is not easy to assess whether the observed influences that were described can be considered significant for the process representation. To investigate this, we apply statistical testing in this section. Figure <xref ref-type="fig" rid="F7"/> shows the distributions of <inline-formula><mml:math id="M109" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-values when testing the significance of signature values calculated with different OFs against each other.</p>

      <fig id="F7" specific-use="star"><label>Figure 7</label><caption><p id="d2e4618">Significance Test between signature representations of the different objective functions. The red (<inline-formula><mml:math id="M110" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.05</mml:mn></mml:mrow></mml:math></inline-formula>) and black (<inline-formula><mml:math id="M111" display="inline"><mml:mrow><mml:mi>p</mml:mi><mml:mo>=</mml:mo><mml:mn mathvariant="normal">0.1</mml:mn></mml:mrow></mml:math></inline-formula>) lines indicate commonly used significance thresholds. Every signature that is below these values can be considered to have a statistically significant difference between the paired samples of signature value distributions (in other words, the two objective functions in the pair lead to statistically different values for the signature across all included models).</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f07.png"/>

        </fig>

      <p id="d2e4651">All signatures show a large range in <inline-formula><mml:math id="M112" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-values, which implies that there are both objective functions which have similar and different representation for each signature. Some OFs typically behave similar, e.g. the distributions of KGE0.2 and log NSE are often comparable for low flow signatures. Figure <xref ref-type="fig" rid="F7"/> shows that the choice of objective function strongly affects the following signatures: the Total RR, Event RR, <inline-formula><mml:math id="M113" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, <inline-formula><mml:math id="M114" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, LFF, BFI, BFRC, and the Flashiness Index. For the remaining 7 signatures (Low and High Flow Duration, HFF, Mean Half Flow Date, FDC Slope, Variability Index, Rising Limb Density) a change of objective function does not lead to a clear shift in signature representation for most of the combinations of objective functions. Table S3, Sect. S2.11 shows the frequency of objective function pairs reaching <inline-formula><mml:math id="M115" display="inline"><mml:mi>p</mml:mi></mml:math></inline-formula>-values of below 0.10 or 0.05 respectively.</p>
      <p id="d2e4693">By comparing Fig. <xref ref-type="fig" rid="F7"/> to  Fig. <xref ref-type="fig" rid="F5"/> we can conclude that there can be considerable spread in the representation of the signatures that do not show a significant difference. This is because the spread for signatures like Low Flow Duration, Variability Index, or Rising Limb Density is often not consistent, meaning that when comparing the distribution of signature values of two objective functions with each other, there is variation as to which one has higher or lower values. Therefore, the significance often shows no significant impact because the signal is not clear. The significance test indicates that the variation is not driven by the objective function and we speculate that in these cases other influences such as catchment characteristics or models drive the spread seen in the signature representation.</p>
</sec>
<sec id="Ch1.S3.SS5">
  <label>3.5</label><title>What controls simulated signature values?</title>
      <p id="d2e4708">Figure <xref ref-type="fig" rid="F8"/> shows the random forest feature importance for catchment, model, and OF (parameter-set variability is considered as a median value across the 500 parameter ensemble).</p>

      <fig id="F8" specific-use="star"><label>Figure 8</label><caption><p id="d2e4715">Random Forest Feature Importance based on predicting the simulated signature value (left) and signature error (right) using the median over the top 500 performing runs.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f08.png"/>

        </fig>

      <p id="d2e4724">Reading the value and error panels together clarifies the role of each factor. Catchment dominates the absolute signature values which is expected, since the basins were deliberately chosen to be hydroclimatically diverse. However, its importance for the signature error is markedly lower and roughly equal across the eight signatures. Objective function has comparatively limited control over the value, yet its importance rises for the error, most strongly for Total RR, Event RR and <inline-formula><mml:math id="M116" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, and typically exceeds that of model choice. Model structure contributes throughout and gains relevance for low-flow and baseflow signatures (BFI, BFRC, <inline-formula><mml:math id="M117" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>), where together with OF choice it challenges the catchment as the leading predictor. The pattern is consistent across signatures: the catchment largely controls the order of magnitude of the simulated signatures, but the choice of objective function and model structure play a large role in how closely those magnitudes match the observations.</p>
</sec>
<sec id="Ch1.S3.SS6">
  <label>3.6</label><title>Overall Skill of Signature Representation</title>
      <p id="d2e4757">Our results suggest that all objective functions can represent some signatures well while having shortcomings in others. This implies that, when selecting an objective function, there is not a single choice that will ensure realistic values across the spectrum of streamflow signatures we investigated. Yet, we can identify objective functions that generally have good process representation across several signatures by investigating the accumulated errors of all significantly influenced signatures as shown in Fig. <xref ref-type="fig" rid="F7"/>. For this, we use Eq. (<xref ref-type="disp-formula" rid="Ch1.E6"/>) to get one median error value per signature and objective function. Figure <xref ref-type="fig" rid="F9"/> shows this accumulated error for the eight tested objective functions.</p>

      <fig id="F9"><label>Figure 9</label><caption><p id="d2e4768">Cumulative error in significantly influenced signatures per tested objective function. This was calculated using the median across the top 500 parameter sets per model.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f09.png"/>

        </fig>

      <p id="d2e4777">The modeled overall error in signature representation varies strongly depending on the objective function chosen. Considering the entire set of signatures with significant metric influence, the KGE-NP, NSE, and DE show the best overall performance. Conversely, this study identifies the log NSE as having the overall largest errors. This, however, does not immediately imply that certain OFs should not be used as each has its strengths and weaknesses. As shown before, the KGE is the best metric for high flow conditions (small errors for <inline-formula><mml:math id="M118" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) and works well for Total RR, while it has considerable weaknesses for low flow representation (e.g. BFI, BFRC and <inline-formula><mml:math id="M119" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>). And while the log NSE performs decently on low flow conditions (<inline-formula><mml:math id="M120" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) it is considerably worse for most other significant signatures. Similarly, despite the KGE-NP showing the best overall performance, it still comes with weaknesses in <inline-formula><mml:math id="M121" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and FI representation.</p>
      <p id="d2e4825">As a robustness check, we repeated the central signature-error and objective-function ranking analysis for the evaluation period. The evaluation-period results broadly confirm the calibration-period patterns. As expected, signature errors generally increase during evaluation, but the relative strengths and weaknesses of the objective functions remain largely similar. In particular, the ranking of objective functions for the OF-sensitive signatures is comparable between calibration and evaluation, suggesting that the main conclusions are not an artefact of calibration-period performance alone. Detailed evaluation period results are provided in the Supplement (Fig. S33).</p>
</sec>
</sec>
<sec id="Ch1.S4">
  <label>4</label><title>Discussion</title>
      <p id="d2e4837">We structured our discussion into three parts. First, we synthesize the main takeaways for signature representation across models, catchments, and objective functions. We condense our findings into some practical recommendations. Second, we discuss the correlation between different signatures and how they relate to the identified impact of objective function choice. Third, we outline the limitations of our study and potential future work.</p>
<sec id="Ch1.S4.SS1">
  <label>4.1</label><title>Main Takeaways and Practical Recommendations</title>
      <p id="d2e4847">Across 47 model structures and 10 hydro-climatically diverse basins, OF choice exerts a systematic but selective influence on simulated streamflow signatures. A paired significance test indicated that 8 of 15 signatures vary significantly with OF choice (Fig. <xref ref-type="fig" rid="F7"/>), whereas timing descriptors such as MHFD and RLD are largely OF-insensitive. This may in part be due to the selection of OFs, as objective functions that specifically target time offset (e.g Mean Absolute Peak Time Error; MAPTE) are less frequently used <xref ref-type="bibr" rid="bib1.bibx16" id="paren.61"/> than the OFs we did select. Generally, no single OF excels at all tested signatures. Considering the cumulative error across OF-sensitive signatures (Fig. <xref ref-type="fig" rid="F9"/>), KGE-NP yields the lowest overall error, with NSE and DE performing competitively in some basins but less consistently across signatures. The DE and the selected transformed OFs (log NSE, KGE0.2) reproduce low flow metrics better (<inline-formula><mml:math id="M122" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, LFF), but particularly the latter come with considerable tradeoffs in the representation of other signatures such as Total RR, baseflow related signatures, or HF signatures.</p>
      <p id="d2e4868">Our analysis hence contributes to the available information of specific strengths and weaknesses of individual OFs. KGE, for example, most reliably reproduces Q95 <xref ref-type="bibr" rid="bib1.bibx48" id="paren.62"/>, and KGE-NP is one of the best OFs for total runoff ratio and for the baseflow index (BFI), while sharing the best representation of baseflow recession (BFRC) with DE. Log NSE, in comparison, showed few strengths for representing hydrological signatures in our experiments, performing well only on <inline-formula><mml:math id="M123" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. However, DE reproduces <inline-formula><mml:math id="M124" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> comparably while carrying a smaller overall error, which is why we recommend DE for that signature and DE or KGE0.2 for low-flow frequency rather than the transformed NSE. The KGE-Split, NSE, and SHE show a balanced error field with individual differences in cumulative errors (Fig. <xref ref-type="fig" rid="F9"/>); within this group, NSE is the most reliable for Event RR, where it ranks alongside KGE-NP, and KGE-Split is otherwise unremarkable but reproduces the Flashiness Index best of all tested OFs.</p>
      <p id="d2e4898">We also compared the influence of OF choice on signature representation with other influencing factors such as catchment characteristics or model structure choice. Our random-forest analysis showed that catchment characteristics are the dominant control of signature representation, with OF choice typically more influential than model choice (Fig. <xref ref-type="fig" rid="F8"/>). Model choice does, however, gain relevance for signatures connected to low flow or baseflow representation, while OF choice seemed especially relevant for baseflow related signatures.</p>
      <p id="d2e4903">When reducing the impact of the conducted catchment selection by training the random forest on signature errors instead of signature values, the importance of the objective functions and model selection increased considerably for most signatures. This indicates the importance of appropriate objective function and model choice across hydro-climatic differences, which may be an overlooked aspect in large sample studies, potentially indicating a path for more detailed evaluation using signatures and their deviations. The broader analysis also contextualizes for which characteristics objective function choice can be influential, while others are predefined by catchment (MHFD) or a combination of catchment and model (RLD), which is shown in Fig. S34. In summary, our results support the argument for a purpose-based model calibration, that considers specific aspects of the flow regime, rather than defaulting to a familiar single metric <xref ref-type="bibr" rid="bib1.bibx43 bib1.bibx29" id="paren.63"/>.</p>

<table-wrap id="T6"><label>Table 6</label><caption><p id="d2e4913">Recommendations for objective functions that can represent the 15 considered signatures well according to our study results. Where a signature was not significantly affected by OF choice, we report “not significant”.</p></caption><oasis:table frame="topbot"><oasis:tgroup cols="2">
     <oasis:colspec colnum="1" colname="col1" align="left"/>
     <oasis:colspec colnum="2" colname="col2" align="left"/>
     <oasis:thead>
       <oasis:row rowsep="1">
         <oasis:entry colname="col1">Signature</oasis:entry>
         <oasis:entry colname="col2">Best OF</oasis:entry>
       </oasis:row>
     </oasis:thead>
     <oasis:tbody>
       <oasis:row>
         <oasis:entry colname="col1">Total Runoff Ratio</oasis:entry>
         <oasis:entry colname="col2">KGE-NP, KGE</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Event Runoff Ratio</oasis:entry>
         <oasis:entry colname="col2">KGE-NP, NSE</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Mean Half Flow Date</oasis:entry>
         <oasis:entry colname="col2">Not significant</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">FDC Slope</oasis:entry>
         <oasis:entry colname="col2">Not significant</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">5th Flow Percentile (<inline-formula><mml:math id="M125" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col2">DE</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Low-Flow Duration</oasis:entry>
         <oasis:entry colname="col2">Not significant</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Low-Flow Frequency</oasis:entry>
         <oasis:entry colname="col2">DE, KGE0.2</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">95th Flow Percentile (<inline-formula><mml:math id="M126" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>)</oasis:entry>
         <oasis:entry colname="col2">KGE</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">High-Flow Duration</oasis:entry>
         <oasis:entry colname="col2">Not significant</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">High-Flow Frequency</oasis:entry>
         <oasis:entry colname="col2">Not significant</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Baseflow Index (BFI)</oasis:entry>
         <oasis:entry colname="col2">KGE-NP</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Baseflow Recession Coefficient (BFRC)</oasis:entry>
         <oasis:entry colname="col2">DE, KGE-NP</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Rising Limb Density (RLD)</oasis:entry>
         <oasis:entry colname="col2">Not significant</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Flashiness Index (FI)</oasis:entry>
         <oasis:entry colname="col2">KGE-Split</oasis:entry>
       </oasis:row>
       <oasis:row>
         <oasis:entry colname="col1">Variability Index (VI)</oasis:entry>
         <oasis:entry colname="col2">Not significant</oasis:entry>
       </oasis:row>
     </oasis:tbody>
   </oasis:tgroup></oasis:table></table-wrap>

      <p id="d2e5099">For the signatures that were found to be significantly impacted by OF choice, we used our results to develop guidelines for which OF will perform well in reproducing a given signature as shown in Table <xref ref-type="table" rid="T6"/>. We can conclude that most OF lead to the best representation for at least one signature, highlighting the need for appropriate selection. However, if the goal is a balanced representation of the water-balance and variability across hydrologic regimes, KGE-NP is a strong option. If low flows are central (ecological flows, drought assessment), preference should be given to DE or KGE0.2. If peak flows are the priority (flood risk, infrastructure design), KGE generally performs best for <inline-formula><mml:math id="M127" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and related high-flow signatures. Reflecting back on our review of metric choice justification (Fig. <xref ref-type="fig" rid="F1"/>), we strongly recommend to document the trade-offs one is willing to accept explicitly (e.g., low-flow degradation under KGE) and, where possible, to verify unaffected signatures (e.g., MHFD, RLD) since OF changes have limited influence there.</p>
</sec>
<sec id="Ch1.S4.SS2">
  <label>4.2</label><title>Impact of Signature Correlations</title>
      <p id="d2e5125">Signatures are not independent from each other. When investigating correlations among observed signatures (Fig. <xref ref-type="fig" rid="F10"/>) we noticed three clusters with similar patterns. First, a <italic>runoff/high-flow</italic> cluster (Total RR, Event RR, and <inline-formula><mml:math id="M128" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>) shows strong positive association, linking long-term water partitioning to the reaction to events. This aligns with aridity being a primary control on mean runoff <xref ref-type="bibr" rid="bib1.bibx8" id="paren.64"/>.  Second, a <italic>baseflow/storage</italic> cluster, including BFI, BFRC and FI, exhibits tight coupling (higher BFI <inline-formula><mml:math id="M129" display="inline"><mml:mo>↔</mml:mo></mml:math></inline-formula> lower BFRC and FI) and a strong relationship to low-flow frequency, reflecting how sustained baseflow suppresses low-flow events. Similar to the first cluster, this group was also sensitive to objective-function (OF) choice. Third, a <italic>variability/threshold</italic> cluster (FDC slope, HFF/LFF, low-flow duration, VI) captures distributional shape and threshold exceedance. These signatures, derived from the hydrograph, partly encode duplicate information and were mostly OF-insensitive. Several other signatures showed weak or inconsistent ties to these clusters, such as High-Flow Duration, Mean Half-Flow Date, Rising Limb Density.</p>

      <fig id="F10" specific-use="star"><label>Figure 10</label><caption><p id="d2e5163">Correlation test between signatures based on Observations. Please note that the signatures have been reorganized in this plot to highlight the emerging patterns.  Section S2.14, Fig. S35 provides a similar plot for the comparing the observed correlations to the modelled ones.</p></caption>
          <graphic xlink:href="https://hess.copernicus.org/articles/30/4867/2026/hess-30-4867-2026-f10.png"/>

        </fig>

      <p id="d2e5172">These clusters help explain why improving one signature typically also improves the other signatures in the cluster, yet they do not imply that different OFs are needed for each cluster. For example, KGE-NP performed best for many signatures in the first two clusters (e.g., Total/Event RR, BFI).</p>
      <p id="d2e5176">We identified some overlap between signatures that are strongly/weakly affected by OF choice and those that were identified as more/less predictable from hydrologic drivers by <xref ref-type="bibr" rid="bib1.bibx2" id="text.65"/>: some signatures (e.g., FDC slope, durations) tend to be OF-insensitive with low spatial predictability, whereas water-balance and high-flow signatures were more OF-sensitive and spatially predictable. A notable exception is Mean Half-Flow Date, which was well predicted in <xref ref-type="bibr" rid="bib1.bibx2" id="text.66"/> but showed little OF sensitivity in our study suggesting a stronger climatic control.</p>
      <p id="d2e5185">If we consider the structural differences among OFs some of these patterns become clearer. For <italic>bias/mean</italic> terms, KGE's mean component directly constrains volumes, explaining its systematic advantage for Total RR and <inline-formula><mml:math id="M130" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">95</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>. OFs that deviate from this approach (NSE, DE, log NSE; Table <xref ref-type="table" rid="T3"/>) typically perform worse for these metrics <xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx48" id="paren.67"/>. Across objective functions, Total RR tends to be underestimated (Fig. <xref ref-type="fig" rid="F4"/>), reflecting systematic water-balance biases related to excessive storage or evapotranspiration. This pattern underscores the importance of bias terms in objective functions, which directly constrain mean flow and promote a more realistic overall water balance <xref ref-type="bibr" rid="bib1.bibx48" id="paren.68"/>. This underestimation is strongest for low-flow-focused objectives, as also noted by <xref ref-type="bibr" rid="bib1.bibx63" id="text.69"/>. Our results do not support the claim that KGE resolves NSE's low-flow insensitivity or performs well simultaneously for high and low flows <xref ref-type="bibr" rid="bib1.bibx3" id="paren.70"/>; note, however, that <xref ref-type="bibr" rid="bib1.bibx3" id="text.71"/> investigate more basins (albeit in a specific biome), while we test a larger number of models.</p>
      <p id="d2e5222">For <italic>variability</italic> terms, replacing variance ratios with FDC-based components (as done for the KGE-NP) or integrating relative errors along the FDC (as done for the DE) re-targets calibration toward distributional shape and mid/low flows, improving BFI, <inline-formula><mml:math id="M131" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula>, and Event RR in our results (Fig. <xref ref-type="fig" rid="F6"/>).</p>
      <p id="d2e5241">For <italic>correlation</italic> terms, moving from Pearson to Spearman (KGE-NP, SHE) reduces sensitivity to extremes and stabilizes correlation. A direct comparison of KGE and SHE, where the dependence component is the only structural change, suggests improved representation for FDC slope, VI, and baseflow metrics, with similar or slightly worse performance for high-flow signatures <xref ref-type="bibr" rid="bib1.bibx33" id="paren.72"/>. Annual weighting using the KGE-Split yielded mixed, context-dependent benefits and no systematic improvement of low-flow characteristics over baseline KGE in our experiments (Fig. <xref ref-type="fig" rid="F9"/>; <xref ref-type="bibr" rid="bib1.bibx18" id="altparen.73"/>).</p>
      <p id="d2e5255">Overall, each objective function emphasizes the hydrograph component it encodes, and no single OF performs best across all signatures. Within this picture, the variability term emerges as a particularly influential lever: changes to it propagated to more signatures than changes to the correlation term, an effect most pronounced for low- and mid-flow signatures, while bias terms remain essential for water-balance and event signatures. This component-specific behaviour is consistent with evidence that signature performance depends strongly on which aspect of the hydrograph a metric targets <xref ref-type="bibr" rid="bib1.bibx48 bib1.bibx45" id="paren.74"/>, and reinforces that OF selection should be guided by the signatures of interest rather than by a search for a universally optimal metric.</p>
      <p id="d2e5261">Streamflow transformations act as a further handle on signature representation, operating not on a single component but on the flow values entering all three. By compressing the dynamic range, the power and log transformations (KGE0.2, log NSE) shift calibration emphasis toward low and mid flows, which is reflected in their improved <inline-formula><mml:math id="M132" display="inline"><mml:mrow><mml:msub><mml:mi>Q</mml:mi><mml:mn mathvariant="normal">5</mml:mn></mml:msub></mml:mrow></mml:math></inline-formula> and LFF representation at the expense of high-flow and water-balance signatures (Fig. <xref ref-type="fig" rid="F6"/>). The choice and strength of transformation is therefore itself a design decision shaping which signatures a calibration favours, complementing the component-level differences discussed above <xref ref-type="bibr" rid="bib1.bibx58" id="paren.75"/>.</p>
</sec>
<sec id="Ch1.S4.SS3">
  <label>4.3</label><title>Limitations and Further Research</title>
      <p id="d2e5288">Our study aimed to fill a specific research gap in the current understanding of the impact of objective function choice on signature representation; using a larger number of models across diverse hydrological regimes (in comparison to Table <xref ref-type="table" rid="T1"/>). To enable a thorough investigation of the results we necessarily had to impose certain limits in other parts of the experimental design (i.e., having to “balance depth with breadth”, <xref ref-type="bibr" rid="bib1.bibx25" id="text.76"/>).</p>
      <p id="d2e5296">First, compared to most studies in Table <xref ref-type="table" rid="T1"/> we use a limited number of catchments. This keeps computational cost manageable and allows more detailed analysis into individual modelling results than would otherwise be possible. We selected these basins to ensure a hydro-climatic spread, as is common practice in these scenarios (see for example the 12 MOPEX basins that have been used for a large number of studies <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx62 bib1.bibx9" id="paren.77"><named-content content-type="pre">e.g.</named-content></xref>. Compared to existing studies, our selected number of objective functions and signatures is fairly typical, while our number of models is clearly higher. The presented study found relevant differences in signature representation depending on the model structure, which might indicate a limitation in generalising signature influence from single-model studies. Conversely, our broader model sample implies that conclusions drawn from the many-basin, single-model studies in Table <xref ref-type="table" rid="T1"/> may themselves be difficult to transfer to other model structures, given that we find model choice to be a non-negligible, signature-dependent control.</p>
      <p id="d2e5308">Second, we focus on single-objective calibration because of its prevalence in the current literature. When multiple aspects of the flow regime matter simultaneously, multi-objective calibration has been shown to provide considerable benefits <xref ref-type="bibr" rid="bib1.bibx28 bib1.bibx50" id="paren.78"/>. Multi-objective formulations can expose the trade-offs explicitly via a Pareto front and allow practitioners to choose solutions that best satisfy competing goals <xref ref-type="bibr" rid="bib1.bibx67 bib1.bibx15 bib1.bibx43" id="paren.79"/>. If a single composite metric must be used, a transparent weighting that reflects decision priorities can be helpful. <xref ref-type="bibr" rid="bib1.bibx63" id="text.80"/>, for instance, argue for directly incorporating multiple ecological-flow characteristics in the OF to improve the relevance of the model for management applications. Similarly, signature values have directly been used in calibration, either as additional objectives or as constraints, to target process realism <xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx68 bib1.bibx64 bib1.bibx17 bib1.bibx51" id="paren.81"/>, and our component-level insights into individual metric strengths and weaknesses can help identify which signatures or complementary metrics to add when a particular aspect of the flow regime is the target. This can potentially improve baseflow behaviour, low-flow frequency, or water-balance partitioning when those aspects are central. However, incorporating signatures does not guarantee a better hindcast/forecast skill in all settings, and can even degrade time-series fit if misapplied <xref ref-type="bibr" rid="bib1.bibx5" id="paren.82"/>.</p>
      <p id="d2e5327">Third, we note the following three methodological limitations in our study design: (1) The application of the median to calculate representative values within the error metric is an important choice, as this does not capture information on the spread of the signature representation. We accounted for that by visually inspecting the underlying violins, but did not explicitly quantify the spread within the applied metrics. (2) The random-forest importance scores should be read as a relative diagnostic within this specific design rather than a universal attribution of variance. The three predictors differ in cardinality (47 models, 10 catchments, 8 objective functions), which can affect permutation-based importance, and benchmark filtering removes poorly performing combinations beforehand. The scores therefore reflect experiment-specific influence rather than general rankings of hydrological controls. (3) Finally, we did not perform any sensitivity analysis before calibrating the models, and simply calibrated all model parameters. Across all calibration runs, approximately 10 % of parameters reached one of their bounds. We found no indication that these boundary hits were systematically concentrated in a specific catchment, model, or objective function. Still, boundary hits can indicate parameter insensitivity, missing processes, data issues, or overly restrictive parameter ranges. They are therefore difficult to interpret in large multi-model calibration experiments and should be considered as an additional diagnostic of calibration reliability in future benchmarking studies.</p>
      <p id="d2e5331">Going forward, we see four priorities to build on the individual strengths of existing work (Table <xref ref-type="table" rid="T1"/>). First, broaden external validity with a larger, stratified basin sample (including ephemeral and groundwater-dominated systems) and potentially a wider model palette spanning physically-based and machine-learning models. Second, test cluster-aware multi-objective calibration that deliberately pairs complementary OFs, e.g., a water-balance/high-flow metric (KGE or KGE-NP) with a low-flow/storage metric (DE or KGE0.2), optionally adding signature-based components. Third, make component-level experiments less ambiguous by swapping KGE-type components one-by-one in a controlled setup, and test weighted KGE-type variants. Finally, propagate forcing/observation and signature-method uncertainty (e.g., baseflow separation) and explore information-rich objectives that go beyond the bias-variability-correlation trio, such as distributional divergences along the flow-duration curve to link calibration more directly to hydrologic processes <xref ref-type="bibr" rid="bib1.bibx34" id="paren.83"/>. Designs like KGE-NP already go in this direction and offer a principled path to retain low-flow sensitivity without relying solely on log transforms.</p>
</sec>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <label>5</label><title>Conclusions</title>
      <p id="d2e5349">We calibrated 47 conceptual model structures across 10 hydro-climatically diverse basins using 8 objective functions and evaluated how well the top-500 parameter sets for each model simulate 15 streamflow signatures. Three results stand out. First, OF choice exerts a selective influence: 8 of 15 tested signatures are OF-sensitive, while others (e.g. mean half-flow date and rising-limb density, and often event durations) change only marginally with OF choice. Second, no single OF performs best across all signatures; aggregated across OF-sensitive signatures, KGE-NP yields the smallest overall error across signatures. Third, differences are interpretable from OF design: metrics emphasizing mean volumes favour water balance and high flows; FDC-based and rank-based terms improve storage/low-flow behaviour. These patterns translate into simple steps for guiding OF choice. For accurate water balance and high-flow magnitude, KGE (or KGE-NP if distributional shape matters) is effective. When low flows and storage-related signatures are central, DE or KGE0.2 should be preferred.</p>
      <p id="d2e5352">While catchment differences dominate the signature values produced, it is the analysis of signature errors that matters for our purpose, since it is unsurprising that geography largely sets observed signature values. When attributing the error in reproducing each signature, the choice of objective function and model structure rise in importance, typically with OF more influential than model structure. Model structure nonetheless exerts a non-negligible, signature-dependent control on this error, particularly for low-flow and baseflow signatures, implying that findings from single-model studies cannot be assumed to generalise easily. Our findings are still bounded by the selected basins, conceptual structures, signatures, and calibration procedure but our use of 47 different models reduces that ambiguity. Priorities for future work include more systematic testing and/or multi-objective formulations, as well as broader regime and model structure coverage.</p>
      <p id="d2e5355">Multiple studies have indicated that OF choice is important for representing hydrologic signatures. Although process-oriented calibration may ultimately benefit most from multi-objective formulations, current practice still often relies on single-metric objective functions. For guidance on process-based modelling efforts, our study provides insights on the impact and limitations of objective function choice in a diverse selection of catchments, conceptual models and selected signatures. This is both useful as a validation approach for existing studies and guidance for further endeavours. Understanding the trade-offs in metric selection offers a practical path to models that reproduce the signatures that matter for the question at hand.</p>
</sec>

      
      </body>
    <back><notes notes-type="codedataavailability"><title>Code and data availability</title>

      <p id="d2e5362">All code used in this analysis will be provided in a GitHub repository <xref ref-type="bibr" rid="bib1.bibx65" id="paren.84"><named-content content-type="pre"><uri>https://github.com/peterwagener/OF_Signature_Code</uri>,</named-content></xref>. The used data can be downloaded from the CARAVAN repository on Zenodo <xref ref-type="bibr" rid="bib1.bibx66" id="paren.85"><named-content content-type="pre"><ext-link xlink:href="https://doi.org/10.5281/zenodo.21399824" ext-link-type="DOI">10.5281/zenodo.21399824</ext-link>,</named-content></xref>.</p>
  </notes><app-group>
        <supplementary-material position="anchor"><p id="d2e5379">The supplement related to this article is available online at <inline-supplementary-material xlink:href="https://doi.org/10.5194/hess-30-4867-2026-supplement" xlink:title="zip">https://doi.org/10.5194/hess-30-4867-2026-supplement</inline-supplementary-material>.</p></supplementary-material>
        </app-group><notes notes-type="authorcontribution"><title>Author contributions</title>

      <p id="d2e5388">PW: methodology, data curation, formal analysis, investigation, visualization, writing – original draft, writing – review and editing; DS: supervision, conceptualization, methodology, investigation, writing – original draft, writing – review and editing; WJMK: supervision, investigation, writing - review and editing; NS: writing – review and editing, supervision.</p>
  </notes><notes notes-type="competinginterests"><title>Competing interests</title>

      <p id="d2e5394">The contact author has declared that none of the authors has any competing interests.</p>
  </notes><notes notes-type="disclaimer"><title>Disclaimer</title>

      <p id="d2e5400">The statements, findings, conclusions, and recommendations are those of the author(s) and do not necessarily reflect the opinions of NOAA. During the preparation of this manuscript, some of the authors used ChatGPT to improve language and readability. The authors reviewed and edited all AI-assisted text and take full responsibility for the content.  Publisher's note: Copernicus Publications remains neutral with regard to jurisdictional claims made in the text, published maps, institutional affiliations, or any other geographical representation in this paper. The authors bear the ultimate responsibility for providing appropriate place names. Views expressed in the text are those of the authors and do not necessarily reflect the views of the publisher.</p>
  </notes><ack><title>Acknowledgements</title><p id="d2e5409">The authors are grateful to the Center for Information Services and High-Performance Computing (Zentrum für Informationsdienste und Hochleistungsrechnen (ZIH)) at Dresden University of Technology, for providing its facilities for high throughput calculations. The authors express their thanks to the editor, reviewers and community contributors for their constructive feedback.</p></ack><notes notes-type="financialsupport"><title>Financial support</title>

      <p id="d2e5414">This research has in part been supported by the National Oceanic and Atmospheric Administration (grant no. NA22NWS4320003).</p>
  </notes><notes notes-type="reviewstatement"><title>Review statement</title>

      <p id="d2e5420">This paper was edited by Markus Hrachowitz and reviewed by Guillaume Thirel and one anonymous referee.</p>
  </notes><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Addor et al.(2017)Addor, Newman, Mizukami, and Clark</label><mixed-citation>Addor, N., Newman, A. J., Mizukami, N., and Clark, M. P.: The CAMELS data set: catchment attributes and meteorology for large-sample studies, Hydrol. Earth Syst. Sci., 21, 5293–5313, <ext-link xlink:href="https://doi.org/10.5194/hess-21-5293-2017" ext-link-type="DOI">10.5194/hess-21-5293-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Addor et al.(2018)Addor, Nearing, Prieto, Newman, Vine, and Clark</label><mixed-citation>Addor, N., Nearing, G., Prieto, C., Newman, A. J., Vine, N. L., and Clark, M. P.: A Ranking of Hydrological Signatures Based on Their Predictability in Space, Water Resour. Res., 54, 8792–8812, <ext-link xlink:href="https://doi.org/10.1029/2018WR022606" ext-link-type="DOI">10.1029/2018WR022606</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Althoff and Rodrigues(2021)</label><mixed-citation>Althoff, D. and Rodrigues, L. N.: Goodness-of-fit criteria for hydrological models: Model calibration and performance assessment, J. Hydrol., 600, 126674, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2021.126674" ext-link-type="DOI">10.1016/j.jhydrol.2021.126674</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Alvarez-Garreton et al.(2018)Alvarez-Garreton, Mendoza, Boisier, Addor, Galleguillos, Zambrano-Bigiarini, Lara, Puelma, Cortes, Garreaud, McPhee, and Ayala</label><mixed-citation>Alvarez-Garreton, C., Mendoza, P. A., Boisier, J. P., Addor, N., Galleguillos, M., Zambrano-Bigiarini, M., Lara, A., Puelma, C., Cortes, G., Garreaud, R., McPhee, J., and Ayala, A.: The CAMELS-CL dataset: catchment attributes and meteorology for large sample studies – Chile dataset, Hydrol. Earth Syst. Sci., 22, 5817–5846, <ext-link xlink:href="https://doi.org/10.5194/hess-22-5817-2018" ext-link-type="DOI">10.5194/hess-22-5817-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Araya et al.(2023)Araya, Mendoza, Muñoz-Castro, and McPhee</label><mixed-citation>Araya, D., Mendoza, P. A., Muñoz-Castro, E., and McPhee, J.: Towards robust seasonal streamflow forecasts in mountainous catchments: impact of calibration metric selection in hydrological modeling, Hydrol. Earth Syst. Sci., 27, 4385–4408, <ext-link xlink:href="https://doi.org/10.5194/hess-27-4385-2023" ext-link-type="DOI">10.5194/hess-27-4385-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Arsenault et al.(2020)Arsenault, Brissette, Martel, Troin, Lévesque, Davidson-Chaput, Gonzalez, Ameli, and Poulin</label><mixed-citation>Arsenault, R., Brissette, F., Martel, J.-L., Troin, M., Lévesque, G., Davidson-Chaput, J., Gonzalez, M. C., Ameli, A., and Poulin, A.: A comprehensive, multisource database for hydrometeorological modeling of 14,425 North American watersheds, Scientific Data, 7, 243, <ext-link xlink:href="https://doi.org/10.1038/s41597-020-00583-2" ext-link-type="DOI">10.1038/s41597-020-00583-2</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Bennett et al.(2013)Bennett, Croke, Guariso, Guillaume, Hamilton, Jakeman, Marsili-Libelli, Newham, Norton, Perrin, Pierce, Robson, Seppelt, Voinov, Fath, and Andreassian</label><mixed-citation>Bennett, N. D., Croke, B. F., Guariso, G., Guillaume, J. H., Hamilton, S. H., Jakeman, A. J., Marsili-Libelli, S., Newham, L. T., Norton, J. P., Perrin, C., Pierce, S. A., Robson, B., Seppelt, R., Voinov, A. A., Fath, B. D., and Andreassian, V.: Characterising performance of environmental models, Environ. Modell. Softw., 40, 1–20, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2012.09.011" ext-link-type="DOI">10.1016/j.envsoft.2012.09.011</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Berghuijs et al.(2017)Berghuijs, Larsen, Van Emmerik, and Woods</label><mixed-citation>Berghuijs, W. R., Larsen, J. R., Van Emmerik, T. H. M., and Woods, R. A.: A Global Assessment of Runoff Sensitivity to Changes in Precipitation, Potential Evaporation, and Other Factors, Water Resour. Res., 53, 8475–8486, <ext-link xlink:href="https://doi.org/10.1002/2017WR021593" ext-link-type="DOI">10.1002/2017WR021593</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Carrillo et al.(2011)Carrillo, Troch, Sivapalan, Wagener, Harman, and Sawicz</label><mixed-citation>Carrillo, G., Troch, P. A., Sivapalan, M., Wagener, T., Harman, C., and Sawicz, K.: Catchment classification: hydrological analysis of catchment behavior through process-based modeling along a climate gradient, Hydrol. Earth Syst. Sci., 15, 3411–3430, <ext-link xlink:href="https://doi.org/10.5194/hess-15-3411-2011" ext-link-type="DOI">10.5194/hess-15-3411-2011</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Chagas et al.(2020)Chagas, Chaffe, Addor, Fan, Fleischmann, Paiva, and Siqueira</label><mixed-citation>Chagas, V. B. P., Chaffe, P. L. B., Addor, N., Fan, F. M., Fleischmann, A. S., Paiva, R. C. D., and Siqueira, V. A.: CAMELS-BR: hydrometeorological time series and landscape attributes for 897 catchments in Brazil, Earth Syst. Sci. Data, 12, 2075–2096, <ext-link xlink:href="https://doi.org/10.5194/essd-12-2075-2020" ext-link-type="DOI">10.5194/essd-12-2075-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Cinkus et al.(2023)Cinkus, Mazzilli, Jourde, Wunsch, Liesch, Ravbar, Chen, and Goldscheider</label><mixed-citation>Cinkus, G., Mazzilli, N., Jourde, H., Wunsch, A., Liesch, T., Ravbar, N., Chen, Z., and Goldscheider, N.: When best is the enemy of good – critical evaluation of performance criteria in hydrological models, Hydrol. Earth Syst. Sci., 27, 2397–2411, <ext-link xlink:href="https://doi.org/10.5194/hess-27-2397-2023" ext-link-type="DOI">10.5194/hess-27-2397-2023</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Clerc-Schwarzenbach et al.(2024)Clerc-Schwarzenbach, Selleri, Neri, Toth, Van Meerveld, and Seibert</label><mixed-citation>Clerc-Schwarzenbach, F., Selleri, G., Neri, M., Toth, E., van Meerveld, I., and Seibert, J.: Large-sample hydrology – a few camels or a whole caravan?, Hydrol. Earth Syst. Sci., 28, 4219–4237, <ext-link xlink:href="https://doi.org/10.5194/hess-28-4219-2024" ext-link-type="DOI">10.5194/hess-28-4219-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Coxon et al.(2020)Coxon, Addor, Bloomfield, Freer, Fry, Hannaford, Howden, Lane, Lewis, Robinson, Wagener, and Woods</label><mixed-citation>Coxon, G., Addor, N., Bloomfield, J. P., Freer, J., Fry, M., Hannaford, J., Howden, N. J. K., Lane, R., Lewis, M., Robinson, E. L., Wagener, T., and Woods, R.: CAMELS-GB: hydrometeorological time series and landscape attributes for 671 catchments in Great Britain, Earth Syst. Sci. Data, 12, 2459–2483, <ext-link xlink:href="https://doi.org/10.5194/essd-12-2459-2020" ext-link-type="DOI">10.5194/essd-12-2459-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Duan et al.(2006)Duan, Schaake, Andréassian, Franks, Goteti, Gupta, Gusev, Habets, Hall, Hay, Hogue, Huang, Leavesley, Liang, Nasonova, Noilhan, Oudin, Sorooshian, Wagener, and Wood</label><mixed-citation>Duan, Q., Schaake, J., Andréassian, V., Franks, S., Goteti, G., Gupta, H., Gusev, Y., Habets, F., Hall, A., Hay, L., Hogue, T., Huang, M., Leavesley, G., Liang, X., Nasonova, O., Noilhan, J., Oudin, L., Sorooshian, S., Wagener, T., and Wood, E.: Model Parameter Estimation Experiment (MOPEX): An overview of science strategy and major results from the second and third workshops, J. Hydrol., 320, 3–17, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2005.07.031" ext-link-type="DOI">10.1016/j.jhydrol.2005.07.031</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Efstratiadis and Koutsoyiannis(2010)</label><mixed-citation>Efstratiadis, A. and Koutsoyiannis, D.: One decade of multi-objective calibration approaches in hydrological modelling: a review, Hydrolog. Sci. J., 55, 58–78, <ext-link xlink:href="https://doi.org/10.1080/02626660903526292" ext-link-type="DOI">10.1080/02626660903526292</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Ehret and Zehe(2011)</label><mixed-citation>Ehret, U. and Zehe, E.: Series distance – an intuitive metric to quantify hydrograph similarity in terms of occurrence, amplitude and timing of hydrological events, Hydrol. Earth Syst. Sci., 15, 877–896, <ext-link xlink:href="https://doi.org/10.5194/hess-15-877-2011" ext-link-type="DOI">10.5194/hess-15-877-2011</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Euser et al.(2013)Euser, Winsemius, Hrachowitz, Fenicia, Uhlenbrook, and Savenije</label><mixed-citation>Euser, T., Winsemius, H. C., Hrachowitz, M., Fenicia, F., Uhlenbrook, S., and Savenije, H. H. G.: A framework to assess the realism of model structures using hydrological signatures, Hydrol. Earth Syst. Sci., 17, 1893–1912, <ext-link xlink:href="https://doi.org/10.5194/hess-17-1893-2013" ext-link-type="DOI">10.5194/hess-17-1893-2013</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Fowler et al.(2018)Fowler, Peel, Western, and Zhang</label><mixed-citation>Fowler, K., Peel, M., Western, A., and Zhang, L.: Improved Rainfall‐Runoff Calibration for Drying Climate: Choice of Objective Function, Water Resour. Res., 54, 3392–3408, <ext-link xlink:href="https://doi.org/10.1029/2017WR022466" ext-link-type="DOI">10.1029/2017WR022466</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Fowler et al.(2021)Fowler, Acharya, Addor, Chou, and Peel</label><mixed-citation>Fowler, K. J. A., Acharya, S. C., Addor, N., Chou, C., and Peel, M. C.: CAMELS-AUS: hydrometeorological time series and landscape attributes for 222 catchments in Australia, Earth Syst. Sci. Data, 13, 3847–3867, <ext-link xlink:href="https://doi.org/10.5194/essd-13-3847-2021" ext-link-type="DOI">10.5194/essd-13-3847-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Färber et al.(2025)Färber, Plessow, Mischel, Kratzert, Addor, Shalev, and Looser</label><mixed-citation>Färber, C., Plessow, H., Mischel, S. A., Kratzert, F., Addor, N., Shalev, G., and Looser, U.: GRDC-Caravan: extending Caravan with data from the Global Runoff Data Centre, Earth Syst. Sci. Data, 17, 4613–4625, <ext-link xlink:href="https://doi.org/10.5194/essd-17-4613-2025" ext-link-type="DOI">10.5194/essd-17-4613-2025</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Garcia et al.(2017)Garcia, Folton, and Oudin</label><mixed-citation>Garcia, F., Folton, N., and Oudin, L.: Which objective function to calibrate rainfall–runoff models for low-flow index simulations?, Hydrolog. Sci. J., 62, 1149–1166, <ext-link xlink:href="https://doi.org/10.1080/02626667.2017.1308511" ext-link-type="DOI">10.1080/02626667.2017.1308511</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Gnann et al.(2021)Gnann, Coxon, Woods, Howden, and McMillan</label><mixed-citation>Gnann, S. J., Coxon, G., Woods, R. A., Howden, N. J., and McMillan, H. K.: TOSSH: A Toolbox for Streamflow Signatures in Hydrology, Environ. Modell. Softw., 138, 104983, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2021.104983" ext-link-type="DOI">10.1016/j.envsoft.2021.104983</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Gupta et al.(2008)Gupta, Wagener, and Liu</label><mixed-citation>Gupta, H. V., Wagener, T., and Liu, Y.: Reconciling theory with observations: elements of a diagnostic approach to model evaluation, Hydrol. Process., 22, 3802–3813, <ext-link xlink:href="https://doi.org/10.1002/hyp.6989" ext-link-type="DOI">10.1002/hyp.6989</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Gupta et al.(2009)Gupta, Kling, Yilmaz, and Martinez</label><mixed-citation>Gupta, H. V., Kling, H., Yilmaz, K. K., and Martinez, G. F.: Decomposition of the mean squared error and NSE performance criteria: Implications for improving hydrological modelling, J. Hydrol., 377, 80–91, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2009.08.003" ext-link-type="DOI">10.1016/j.jhydrol.2009.08.003</ext-link>, 2009.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Gupta et al.(2014)Gupta, Perrin, Blöschl, Montanari, Kumar, Clark, and Andréassian</label><mixed-citation>Gupta, H. V., Perrin, C., Blöschl, G., Montanari, A., Kumar, R., Clark, M., and Andréassian, V.: Large-sample hydrology: a need to balance depth with breadth, Hydrol. Earth Syst. Sci., 18, 463–477, <ext-link xlink:href="https://doi.org/10.5194/hess-18-463-2014" ext-link-type="DOI">10.5194/hess-18-463-2014</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Hallouin et al.(2020)</label><mixed-citation>Hallouin, T., Bruen, M., and O'Loughlin, F. E.: Calibration of hydrological models for ecologically relevant streamflow predictions: a trade-off between fitting well to data and estimating consistent parameter sets?, Hydrol. Earth Syst. Sci., 24, 1031–1054, <ext-link xlink:href="https://doi.org/10.5194/hess-24-1031-2020" ext-link-type="DOI">10.5194/hess-24-1031-2020</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Hansen et al.(2003)Hansen, Müller, and Koumoutsakos</label><mixed-citation>Hansen, N., Müller, S. D., and Koumoutsakos, P.: Reducing the Time Complexity of the Derandomized Evolution Strategy with Covariance Matrix Adaptation (CMA-ES), Evol. Comput., 11, 1–18, <ext-link xlink:href="https://doi.org/10.1162/106365603321828970" ext-link-type="DOI">10.1162/106365603321828970</ext-link>, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Hernandez-Suarez et al.(2018)Hernandez-Suarez, Nejadhashemi, Kropp, Abouali, Zhang, and Deb</label><mixed-citation>Hernandez-Suarez, J. S., Nejadhashemi, A. P., Kropp, I. M., Abouali, M., Zhang, Z., and Deb, K.: Evaluation of the impacts of hydrologic model calibration methods on predictability of ecologically-relevant hydrologic indices, J. Hydrol., 564, 758–772, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2018.07.056" ext-link-type="DOI">10.1016/j.jhydrol.2018.07.056</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx29"><label>Jackson et al.(2019)Jackson, Roberts, Nelsen, Williams, Nelson, and Ames</label><mixed-citation>Jackson, E. K., Roberts, W., Nelsen, B., Williams, G. P., Nelson, E. J., and Ames, D. P.: Introductory overview: Error metrics for hydrologic modelling – A review of common practices and an open source library to facilitate use and adoption, Environ. Modell. Softw., 119, 32–48, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2019.05.001" ext-link-type="DOI">10.1016/j.envsoft.2019.05.001</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>Jacobs et al.(2024)Jacobs, Tobi, and Hengeveld</label><mixed-citation>Jacobs, B., Tobi, H., and Hengeveld, G. M.: Linking error measures to model questions, Ecol. Model., 487, 110562, <ext-link xlink:href="https://doi.org/10.1016/j.ecolmodel.2023.110562" ext-link-type="DOI">10.1016/j.ecolmodel.2023.110562</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Jakeman et al.(2006)Jakeman, Letcher, and Norton</label><mixed-citation>Jakeman, A., Letcher, R., and Norton, J.: Ten iterative steps in development and evaluation of environmental models, Environ. Modell. Softw., 21, 602–614, <ext-link xlink:href="https://doi.org/10.1016/j.envsoft.2006.01.004" ext-link-type="DOI">10.1016/j.envsoft.2006.01.004</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Jakeman et al.(2018)Jakeman, Sawah, Cuddy, Robson, McIntyre, and Cook</label><mixed-citation>Jakeman, A. J., Sawah, S. E., Cuddy, S., Robson, B., McIntyre, N., and Cook, F.: Good Modelling Practice Principles, Tech. rep., The State of Queensland (Department of Environment and Science, <uri>https://science.desi.qld.gov.au/__data/assets/pdf_file/0011/81110/qwmn-good-modelling-practice-principles.pdf</uri> (last access: 23 July 2026), 2018.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Kiraz et al.(2023)Kiraz, Coxon, and Wagener</label><mixed-citation>Kiraz, M., Coxon, G., and Wagener, T.: A Signature‐Based Hydrologic Efficiency Metric for Model Calibration and Evaluation in Gauged and Ungauged Catchments, Water Resour. Res., 59, <ext-link xlink:href="https://doi.org/10.1029/2023WR035321" ext-link-type="DOI">10.1029/2023WR035321</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx34"><label>Kirchner(2006)</label><mixed-citation>Kirchner, J. W.: Getting the right answers for the right reasons: Linking measurements, analyses, and models to advance the science of hydrology, Water Resour. Res., 42, <ext-link xlink:href="https://doi.org/10.1029/2005WR004362" ext-link-type="DOI">10.1029/2005WR004362</ext-link>, 2006.</mixed-citation></ref>
      <ref id="bib1.bibx35"><label>Klingler et al.(2021)Klingler, Schulz, and Herrnegger</label><mixed-citation>Klingler, C., Schulz, K., and Herrnegger, M.: LamaH-CE: LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe, Earth Syst. Sci. Data, 13, 4529–4565, <ext-link xlink:href="https://doi.org/10.5194/essd-13-4529-2021" ext-link-type="DOI">10.5194/essd-13-4529-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx36"><label>Knoben(2024)</label><mixed-citation>Knoben, W. J. M.: Setting expectations for hydrologic model performance with an ensemble of simple benchmarks, Hydrol. Process., 38, <ext-link xlink:href="https://doi.org/10.1002/hyp.15288" ext-link-type="DOI">10.1002/hyp.15288</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx37"><label>Knoben et al.(2018)Knoben, Woods, and Freer</label><mixed-citation>Knoben, W. J. M., Woods, R. A., and Freer, J. E.: A Quantitative Hydrological Climate Classification Evaluated With Independent Streamflow Data, Water Resour. Res., 54, 5088–5109, <ext-link xlink:href="https://doi.org/10.1029/2018WR022913" ext-link-type="DOI">10.1029/2018WR022913</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx38"><label>Knoben et al.(2019a)Knoben, Freer, Fowler, Peel, and Woods</label><mixed-citation>Knoben, W. J. M., Freer, J. E., Fowler, K. J. A., Peel, M. C., and Woods, R. A.: Modular Assessment of Rainfall–Runoff Models Toolbox (MARRMoT) v1.2: an open-source, extendable framework providing implementations of 46 conceptual hydrologic models as continuous state-space formulations, Geosci. Model Dev., 12, 2463–2480, <ext-link xlink:href="https://doi.org/10.5194/gmd-12-2463-2019" ext-link-type="DOI">10.5194/gmd-12-2463-2019</ext-link>, 2019a.</mixed-citation></ref>
      <ref id="bib1.bibx39"><label>Knoben et al.(2019b)Knoben, Freer, and Woods</label><mixed-citation>Knoben, W. J. M., Freer, J. E., and Woods, R. A.: Technical note: Inherent benchmark or not? Comparing Nash–Sutcliffe and Kling–Gupta efficiency scores, Hydrol. Earth Syst. Sci., 23, 4323–4331, <ext-link xlink:href="https://doi.org/10.5194/hess-23-4323-2019" ext-link-type="DOI">10.5194/hess-23-4323-2019</ext-link>, 2019b.</mixed-citation></ref>
      <ref id="bib1.bibx40"><label>Knoben et al.(2020)Knoben, Freer, Peel, Fowler, and Woods</label><mixed-citation>Knoben, W. J. M., Freer, J. E., Peel, M. C., Fowler, K. J. A., and Woods, R. A.: A Brief Analysis of Conceptual Model Structure Uncertainty Using 36 Models and 559 Catchments, Water Resour. Res., 56, <ext-link xlink:href="https://doi.org/10.1029/2019wr025975" ext-link-type="DOI">10.1029/2019wr025975</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx41"><label>Kratzert et al.(2023)Kratzert, Nearing, Addor, Erickson, Gauch, Gilon, Gudmundsson, Hassidim, Klotz, Nevo, Shalev, and Matias</label><mixed-citation>Kratzert, F., Nearing, G., Addor, N., Erickson, T., Gauch, M., Gilon, O., Gudmundsson, L., Hassidim, A., Klotz, D., Nevo, S., Shalev, G., and Matias, Y.: Caravan – A global community dataset for large-sample hydrology, Scientific Data, 10, <ext-link xlink:href="https://doi.org/10.1038/s41597-023-01975-w" ext-link-type="DOI">10.1038/s41597-023-01975-w</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx42"><label>Krause et al.(2005)Krause, Boyle, and Bäse</label><mixed-citation>Krause, P., Boyle, D. P., and Bäse, F.: Comparison of different efficiency criteria for hydrological model assessment, Adv. Geosci., 5, 89–97, <ext-link xlink:href="https://doi.org/10.5194/adgeo-5-89-2005" ext-link-type="DOI">10.5194/adgeo-5-89-2005</ext-link>, 2005.</mixed-citation></ref>
      <ref id="bib1.bibx43"><label>Mai(2023)</label><mixed-citation>Mai, J.: Ten strategies towards successful calibration of environmental models, J. Hydrol., 620, 129414, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2023.129414" ext-link-type="DOI">10.1016/j.jhydrol.2023.129414</ext-link>, 2023.</mixed-citation></ref>
      <ref id="bib1.bibx44"><label>McMillan(2020)</label><mixed-citation>McMillan, H.: Linking hydrologic signatures to hydrologic processes: A review, Hydrol. Process., 34, 1393–1409, <ext-link xlink:href="https://doi.org/10.1002/hyp.13632" ext-link-type="DOI">10.1002/hyp.13632</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx45"><label>Melsen et al.(2019)Melsen, Teuling, Torfs, Zappa, Mizukami, Mendoza, Clark, and Uijlenhoet</label><mixed-citation>Melsen, L. A., Teuling, A. J., Torfs, P. J., Zappa, M., Mizukami, N., Mendoza, P. A., Clark, M. P., and Uijlenhoet, R.: Subjective modeling decisions can significantly impact the simulation of flood and drought events, J. Hydrol., 568, 1093–1104, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2018.11.046" ext-link-type="DOI">10.1016/j.jhydrol.2018.11.046</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx46"><label>Melsen et al.(2025)Melsen, Puy, Torfs, and Saltelli</label><mixed-citation>Melsen, L. A., Puy, A., Torfs, P. J. J. F., and Saltelli, A.: The rise of the Nash-Sutcliffe efficiency in hydrology, Hydrolog. Sci. J., 70, 1248–1259, <ext-link xlink:href="https://doi.org/10.1080/02626667.2025.2475105" ext-link-type="DOI">10.1080/02626667.2025.2475105</ext-link>, 2025.</mixed-citation></ref>
      <ref id="bib1.bibx47"><label>Mendoza et al.(2016)Mendoza, Clark, Mizukami, Gutmann, Arnold, Brekke, and Rajagopalan</label><mixed-citation>Mendoza, P. A., Clark, M. P., Mizukami, N., Gutmann, E. D., Arnold, J. R., Brekke, L. D., and Rajagopalan, B.: How do hydrologic modeling decisions affect the portrayal of climate change impacts?, Hydrol. Process., 30, 1071–1095, <ext-link xlink:href="https://doi.org/10.1002/hyp.10684" ext-link-type="DOI">10.1002/hyp.10684</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx48"><label>Mizukami et al.(2019)Mizukami, Rakovec, Newman, Clark, Wood, Gupta, and Kumar</label><mixed-citation>Mizukami, N., Rakovec, O., Newman, A. J., Clark, M. P., Wood, A. W., Gupta, H. V., and Kumar, R.: On the choice of calibration metrics for “high-flow” estimation using hydrologic models, Hydrol. Earth Syst. Sci., 23, 2601–2614, <ext-link xlink:href="https://doi.org/10.5194/hess-23-2601-2019" ext-link-type="DOI">10.5194/hess-23-2601-2019</ext-link>, 2019.</mixed-citation></ref>
      <ref id="bib1.bibx49"><label>Nash and Sutcliffe(1970)</label><mixed-citation>Nash, J. and Sutcliffe, J.: River flow forecasting through conceptual models part I – A discussion of principles, J. Hydrol., 10, 282–290, <ext-link xlink:href="https://doi.org/10.1016/0022-1694(70)90255-6" ext-link-type="DOI">10.1016/0022-1694(70)90255-6</ext-link>, 1970.</mixed-citation></ref>
      <ref id="bib1.bibx50"><label>Nemri and Kinnard(2020)</label><mixed-citation>Nemri, S. and Kinnard, C.: Comparing calibration strategies of a conceptual snow hydrology model and their impact on model performance and parameter identifiability, J. Hydrol., 582, 124474, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2019.124474" ext-link-type="DOI">10.1016/j.jhydrol.2019.124474</ext-link>, 2020.</mixed-citation></ref>
      <ref id="bib1.bibx51"><label>Pool et al.(2017)Pool, Vis, Knight, and Seibert</label><mixed-citation>Pool, S., Vis, M. J. P., Knight, R. R., and Seibert, J.: Streamflow characteristics from modeled runoff time series – importance of calibration criteria selection, Hydrol. Earth Syst. Sci., 21, 5443–5457, <ext-link xlink:href="https://doi.org/10.5194/hess-21-5443-2017" ext-link-type="DOI">10.5194/hess-21-5443-2017</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx52"><label>Pool et al.(2018)Pool, Vis, and Seibert</label><mixed-citation>Pool, S., Vis, M., and Seibert, J.: Evaluating model performance: towards a non-parametric variant of the Kling-Gupta efficiency, Hydrolog. Sci. J., 63, 1941–1953, <ext-link xlink:href="https://doi.org/10.1080/02626667.2018.1552002" ext-link-type="DOI">10.1080/02626667.2018.1552002</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx53"><label>Santos et al.(2018)Santos, Thirel, and Perrin</label><mixed-citation>Santos, L., Thirel, G., and Perrin, C.: Technical note: Pitfalls in using log-transformed flows within the KGE criterion, Hydrol. Earth Syst. Sci., 22, 4583–4591, <ext-link xlink:href="https://doi.org/10.5194/hess-22-4583-2018" ext-link-type="DOI">10.5194/hess-22-4583-2018</ext-link>, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx54"><label>Schaefli and Gupta(2007)</label><mixed-citation>Schaefli, B. and Gupta, H. V.: Do Nash values have value?, Hydrol. Process., 21, 2075–2080, <ext-link xlink:href="https://doi.org/10.1002/hyp.6825" ext-link-type="DOI">10.1002/hyp.6825</ext-link>, 2007.</mixed-citation></ref>
      <ref id="bib1.bibx55"><label>Schwemmle et al.(2021)Schwemmle, Demand, and Weiler</label><mixed-citation>Schwemmle, R., Demand, D., and Weiler, M.: Technical note: Diagnostic efficiency – specific evaluation of model performance, Hydrol. Earth Syst. Sci., 25, 2187–2198, <ext-link xlink:href="https://doi.org/10.5194/hess-25-2187-2021" ext-link-type="DOI">10.5194/hess-25-2187-2021</ext-link>, 2021.</mixed-citation></ref>
      <ref id="bib1.bibx56"><label>Seibert(2001)</label><mixed-citation>Seibert, J.: On the need for benchmarks in hydrological modelling, Hydrol. Process., 15, 1063–1064, <ext-link xlink:href="https://doi.org/10.1002/hyp.446" ext-link-type="DOI">10.1002/hyp.446</ext-link>, 2001.</mixed-citation></ref>
      <ref id="bib1.bibx57"><label>Seiller et al.(2017)Seiller, Roy, and Anctil</label><mixed-citation>Seiller, G., Roy, R., and Anctil, F.: Influence of three common calibration metrics on the diagnosis of climate change impacts on water resources, J. Hydrol., 547, 280–295, <ext-link xlink:href="https://doi.org/10.1016/j.jhydrol.2017.02.004" ext-link-type="DOI">10.1016/j.jhydrol.2017.02.004</ext-link>, 2017. </mixed-citation></ref>
      <ref id="bib1.bibx58"><label>Thirel et al.(2024)Thirel, Santos, Delaigue, and Perrin</label><mixed-citation>Thirel, G., Santos, L., Delaigue, O., and Perrin, C.: On the use of streamflow transformations for hydrological model calibration, Hydrol. Earth Syst. Sci., 28, 4837–4860, <ext-link xlink:href="https://doi.org/10.5194/hess-28-4837-2024" ext-link-type="DOI">10.5194/hess-28-4837-2024</ext-link>, 2024.</mixed-citation></ref>
      <ref id="bib1.bibx59"><label>Trotter and Knoben(2022)</label><mixed-citation>Trotter, L. and Knoben, W. J. M.: MARRMoT v2.1, Zenodo [code], <ext-link xlink:href="https://doi.org/10.5281/zenodo.6484372" ext-link-type="DOI">10.5281/zenodo.6484372</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx60"><label>Trotter et al.(2022)Trotter, Knoben, Fowler, Saft, and Peel</label><mixed-citation>Trotter, L., Knoben, W. J. M., Fowler, K. J. A., Saft, M., and Peel, M. C.: Modular Assessment of Rainfall–Runoff Models Toolbox (MARRMoT) v2.1: an object-oriented implementation of 47 established hydrological models for improved speed and readability, Geosci. Model Dev., 15, 6359–6369, <ext-link xlink:href="https://doi.org/10.5194/gmd-15-6359-2022" ext-link-type="DOI">10.5194/gmd-15-6359-2022</ext-link>, 2022.</mixed-citation></ref>
      <ref id="bib1.bibx61"><label>van Waveren et al.(1999)Waveren, Groot, Scholten, Geer, Wösten, Koeze, and Noort</label><mixed-citation> van Waveren, R. H., Groot, S., Scholten, H., Geer, F. v., Wösten, J., Koeze, R., and Noort, J.: Good Modelling Practice Handbook, Tech. rep., Dutch Dept. of Public Works, Institute for Inland Water Management and Waste Water Treatment, report 99.036, ISBN 9057730561, 1999.</mixed-citation></ref>
      <ref id="bib1.bibx62"><label>van Werkhoven et al.(2008)van Werkhoven, Wagener, Reed, and Tang</label><mixed-citation>van Werkhoven, K., Wagener, T., Reed, P., and Tang, Y.: Characterization of watershed model behavior across a hydroclimatic gradient, Water Resour. Res., 44, <ext-link xlink:href="https://doi.org/10.1029/2007WR006271" ext-link-type="DOI">10.1029/2007WR006271</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx63"><label>Vis et al.(2015)Vis, Knight, Pool, Wolfe, and Seibert</label><mixed-citation>Vis, M., Knight, R., Pool, S., Wolfe, W., and Seibert, J.: Model Calibration Criteria for Estimating Ecological Flow Characteristics, Water, 7, 2358–2381, <ext-link xlink:href="https://doi.org/10.3390/w7052358" ext-link-type="DOI">10.3390/w7052358</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx64"><label>Wagener and Montanari(2011)</label><mixed-citation>Wagener, T. and Montanari, A.: Convergence of approaches toward reducing uncertainty in predictions in ungauged basins, Water Resour. Res., 47, 2010WR009469, <ext-link xlink:href="https://doi.org/10.1029/2010WR009469" ext-link-type="DOI">10.1029/2010WR009469</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx65"><label>Wagener(2025)</label><mixed-citation>Wagener, P.: Code for Analysis and Visualization, GitHub [code], <uri>https://github.com/peterwagener/OF_Signature_Code.git</uri> (last access: 16 July 2026), 2025.</mixed-citation></ref>
      <ref id="bib1.bibx66"><label>Wagener(2026)</label><mixed-citation>Wagener, P.: Calibration and Evaluation Data for Objective Function Impact on Hydrologic Signatures, Zenodo [data set], <ext-link xlink:href="https://doi.org/10.5281/zenodo.21399824" ext-link-type="DOI">10.5281/zenodo.21399824</ext-link>, 2026.</mixed-citation></ref>
      <ref id="bib1.bibx67"><label>Yapo et al.(1998)Yapo, Gupta, and Sorooshian</label><mixed-citation>Yapo, P. O., Gupta, H. V., and Sorooshian, S.: Multi-objective global optimization for hydrologic models, J. Hydrol., 204, 83–97, <ext-link xlink:href="https://doi.org/10.1016/S0022-1694(97)00107-8" ext-link-type="DOI">10.1016/S0022-1694(97)00107-8</ext-link>, 1998.</mixed-citation></ref>
      <ref id="bib1.bibx68"><label>Yilmaz et al.(2008)Yilmaz, Gupta, and Wagener</label><mixed-citation>Yilmaz, K. K., Gupta, H. V., and Wagener, T.: A process‐based diagnostic approach to model evaluation: Application to the NWS distributed hydrologic model, Water Resour. Res., 44, 2007WR006716, <ext-link xlink:href="https://doi.org/10.1029/2007WR006716" ext-link-type="DOI">10.1029/2007WR006716</ext-link>, 2008.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Metrics that matter: objective functions and their impact on signature representation in conceptual hydrological models</article-title-html>
<abstract-html/>
<ref-html id="bib1.bib1"><label>Addor et al.(2017)Addor, Newman, Mizukami, and
Clark</label><mixed-citation>
      
Addor, N., Newman, A. J., Mizukami, N., and Clark, M. P.: The CAMELS data set: catchment attributes and meteorology for large-sample studies, Hydrol. Earth Syst. Sci., 21, 5293–5313, <a href="https://doi.org/10.5194/hess-21-5293-2017" target="_blank">https://doi.org/10.5194/hess-21-5293-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Addor et al.(2018)Addor, Nearing, Prieto, Newman, Vine, and
Clark</label><mixed-citation>
      
Addor, N., Nearing, G., Prieto, C., Newman, A. J., Vine, N. L., and Clark,
M. P.: A Ranking of Hydrological Signatures Based on Their Predictability in
Space, Water Resour. Res., 54, 8792–8812, <a href="https://doi.org/10.1029/2018WR022606" target="_blank">https://doi.org/10.1029/2018WR022606</a>,
2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Althoff and Rodrigues(2021)</label><mixed-citation>
      
Althoff, D. and Rodrigues, L. N.: Goodness-of-fit criteria for hydrological
models: Model calibration and performance assessment, J. Hydrol.,
600, 126674, <a href="https://doi.org/10.1016/j.jhydrol.2021.126674" target="_blank">https://doi.org/10.1016/j.jhydrol.2021.126674</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Alvarez-Garreton et al.(2018)Alvarez-Garreton, Mendoza, Boisier,
Addor, Galleguillos, Zambrano-Bigiarini, Lara, Puelma, Cortes, Garreaud,
McPhee, and Ayala</label><mixed-citation>
      
Alvarez-Garreton, C., Mendoza, P. A., Boisier, J. P., Addor, N., Galleguillos, M., Zambrano-Bigiarini, M., Lara, A., Puelma, C., Cortes, G., Garreaud, R., McPhee, J., and Ayala, A.: The CAMELS-CL dataset: catchment attributes and meteorology for large sample studies – Chile dataset, Hydrol. Earth Syst. Sci., 22, 5817–5846, <a href="https://doi.org/10.5194/hess-22-5817-2018" target="_blank">https://doi.org/10.5194/hess-22-5817-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Araya et al.(2023)Araya, Mendoza, Muñoz-Castro, and
McPhee</label><mixed-citation>
      
Araya, D., Mendoza, P. A., Muñoz-Castro, E., and McPhee, J.: Towards robust seasonal streamflow forecasts in mountainous catchments: impact of calibration metric selection in hydrological modeling, Hydrol. Earth Syst. Sci., 27, 4385–4408, <a href="https://doi.org/10.5194/hess-27-4385-2023" target="_blank">https://doi.org/10.5194/hess-27-4385-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Arsenault et al.(2020)Arsenault, Brissette, Martel, Troin, Lévesque,
Davidson-Chaput, Gonzalez, Ameli, and
Poulin</label><mixed-citation>
      
Arsenault, R., Brissette, F., Martel, J.-L., Troin, M., Lévesque, G.,
Davidson-Chaput, J., Gonzalez, M. C., Ameli, A., and Poulin, A.: A
comprehensive, multisource database for hydrometeorological modeling of
14,425 North American watersheds, Scientific Data, 7, 243,
<a href="https://doi.org/10.1038/s41597-020-00583-2" target="_blank">https://doi.org/10.1038/s41597-020-00583-2</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Bennett et al.(2013)Bennett, Croke, Guariso, Guillaume, Hamilton,
Jakeman, Marsili-Libelli, Newham, Norton, Perrin, Pierce, Robson, Seppelt,
Voinov, Fath, and Andreassian</label><mixed-citation>
      
Bennett, N. D., Croke, B. F., Guariso, G., Guillaume, J. H., Hamilton, S. H.,
Jakeman, A. J., Marsili-Libelli, S., Newham, L. T., Norton, J. P., Perrin,
C., Pierce, S. A., Robson, B., Seppelt, R., Voinov, A. A., Fath, B. D., and
Andreassian, V.: Characterising performance of environmental models,
Environ. Modell. Softw., 40, 1–20,
<a href="https://doi.org/10.1016/j.envsoft.2012.09.011" target="_blank">https://doi.org/10.1016/j.envsoft.2012.09.011</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Berghuijs et al.(2017)Berghuijs, Larsen, Van Emmerik, and
Woods</label><mixed-citation>
      
Berghuijs, W. R., Larsen, J. R., Van Emmerik, T. H. M., and Woods, R. A.: A
Global Assessment of Runoff Sensitivity to Changes in
Precipitation, Potential Evaporation, and Other Factors, Water
Resour. Res., 53, 8475–8486, <a href="https://doi.org/10.1002/2017WR021593" target="_blank">https://doi.org/10.1002/2017WR021593</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Carrillo et al.(2011)Carrillo, Troch, Sivapalan, Wagener, Harman, and
Sawicz</label><mixed-citation>
      
Carrillo, G., Troch, P. A., Sivapalan, M., Wagener, T., Harman, C., and Sawicz, K.: Catchment classification: hydrological analysis of catchment behavior through process-based modeling along a climate gradient, Hydrol. Earth Syst. Sci., 15, 3411–3430, <a href="https://doi.org/10.5194/hess-15-3411-2011" target="_blank">https://doi.org/10.5194/hess-15-3411-2011</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Chagas et al.(2020)Chagas, Chaffe, Addor, Fan, Fleischmann, Paiva,
and Siqueira</label><mixed-citation>
      
Chagas, V. B. P., Chaffe, P. L. B., Addor, N., Fan, F. M., Fleischmann, A. S., Paiva, R. C. D., and Siqueira, V. A.: CAMELS-BR: hydrometeorological time series and landscape attributes for 897 catchments in Brazil, Earth Syst. Sci. Data, 12, 2075–2096, <a href="https://doi.org/10.5194/essd-12-2075-2020" target="_blank">https://doi.org/10.5194/essd-12-2075-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Cinkus et al.(2023)Cinkus, Mazzilli, Jourde, Wunsch, Liesch, Ravbar,
Chen, and Goldscheider</label><mixed-citation>
      
Cinkus, G., Mazzilli, N., Jourde, H., Wunsch, A., Liesch, T., Ravbar, N., Chen, Z., and Goldscheider, N.: When best is the enemy of good – critical evaluation of performance criteria in hydrological models, Hydrol. Earth Syst. Sci., 27, 2397–2411, <a href="https://doi.org/10.5194/hess-27-2397-2023" target="_blank">https://doi.org/10.5194/hess-27-2397-2023</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Clerc-Schwarzenbach et al.(2024)Clerc-Schwarzenbach, Selleri, Neri,
Toth, Van Meerveld, and
Seibert</label><mixed-citation>
      
Clerc-Schwarzenbach, F., Selleri, G., Neri, M., Toth, E., van Meerveld, I., and Seibert, J.: Large-sample hydrology – a few camels or a whole caravan?, Hydrol. Earth Syst. Sci., 28, 4219–4237, <a href="https://doi.org/10.5194/hess-28-4219-2024" target="_blank">https://doi.org/10.5194/hess-28-4219-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Coxon et al.(2020)Coxon, Addor, Bloomfield, Freer, Fry, Hannaford,
Howden, Lane, Lewis, Robinson, Wagener, and
Woods</label><mixed-citation>
      
Coxon, G., Addor, N., Bloomfield, J. P., Freer, J., Fry, M., Hannaford, J., Howden, N. J. K., Lane, R., Lewis, M., Robinson, E. L., Wagener, T., and Woods, R.: CAMELS-GB: hydrometeorological time series and landscape attributes for 671 catchments in Great Britain, Earth Syst. Sci. Data, 12, 2459–2483, <a href="https://doi.org/10.5194/essd-12-2459-2020" target="_blank">https://doi.org/10.5194/essd-12-2459-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Duan et al.(2006)Duan, Schaake, Andréassian, Franks, Goteti, Gupta,
Gusev, Habets, Hall, Hay, Hogue, Huang, Leavesley, Liang, Nasonova, Noilhan,
Oudin, Sorooshian, Wagener, and Wood</label><mixed-citation>
      
Duan, Q., Schaake, J., Andréassian, V., Franks, S., Goteti, G., Gupta, H.,
Gusev, Y., Habets, F., Hall, A., Hay, L., Hogue, T., Huang, M., Leavesley,
G., Liang, X., Nasonova, O., Noilhan, J., Oudin, L., Sorooshian, S., Wagener,
T., and Wood, E.: Model Parameter Estimation Experiment (MOPEX): An
overview of science strategy and major results from the second and third
workshops, J. Hydrol., 320, 3–17,
<a href="https://doi.org/10.1016/j.jhydrol.2005.07.031" target="_blank">https://doi.org/10.1016/j.jhydrol.2005.07.031</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Efstratiadis and Koutsoyiannis(2010)</label><mixed-citation>
      
Efstratiadis, A. and Koutsoyiannis, D.: One decade of multi-objective
calibration approaches in hydrological modelling: a review, Hydrolog.
Sci. J., 55, 58–78, <a href="https://doi.org/10.1080/02626660903526292" target="_blank">https://doi.org/10.1080/02626660903526292</a>, 2010.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Ehret and Zehe(2011)</label><mixed-citation>
      
Ehret, U. and Zehe, E.: Series distance – an intuitive metric to quantify hydrograph similarity in terms of occurrence, amplitude and timing of hydrological events, Hydrol. Earth Syst. Sci., 15, 877–896, <a href="https://doi.org/10.5194/hess-15-877-2011" target="_blank">https://doi.org/10.5194/hess-15-877-2011</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Euser et al.(2013)Euser, Winsemius, Hrachowitz, Fenicia, Uhlenbrook,
and Savenije</label><mixed-citation>
      
Euser, T., Winsemius, H. C., Hrachowitz, M., Fenicia, F., Uhlenbrook, S., and Savenije, H. H. G.: A framework to assess the realism of model structures using hydrological signatures, Hydrol. Earth Syst. Sci., 17, 1893–1912, <a href="https://doi.org/10.5194/hess-17-1893-2013" target="_blank">https://doi.org/10.5194/hess-17-1893-2013</a>, 2013.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Fowler et al.(2018)Fowler, Peel, Western, and Zhang</label><mixed-citation>
      
Fowler, K., Peel, M., Western, A., and Zhang, L.: Improved Rainfall‐Runoff
Calibration for Drying Climate: Choice of Objective Function, Water Resour.
Res., 54, 3392–3408, <a href="https://doi.org/10.1029/2017WR022466" target="_blank">https://doi.org/10.1029/2017WR022466</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Fowler et al.(2021)Fowler, Acharya, Addor, Chou, and
Peel</label><mixed-citation>
      
Fowler, K. J. A., Acharya, S. C., Addor, N., Chou, C., and Peel, M. C.: CAMELS-AUS: hydrometeorological time series and landscape attributes for 222 catchments in Australia, Earth Syst. Sci. Data, 13, 3847–3867, <a href="https://doi.org/10.5194/essd-13-3847-2021" target="_blank">https://doi.org/10.5194/essd-13-3847-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Färber et al.(2025)Färber, Plessow, Mischel, Kratzert, Addor,
Shalev, and Looser</label><mixed-citation>
      
Färber, C., Plessow, H., Mischel, S. A., Kratzert, F., Addor, N., Shalev, G., and Looser, U.: GRDC-Caravan: extending Caravan with data from the Global Runoff Data Centre, Earth Syst. Sci. Data, 17, 4613–4625, <a href="https://doi.org/10.5194/essd-17-4613-2025" target="_blank">https://doi.org/10.5194/essd-17-4613-2025</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Garcia et al.(2017)Garcia, Folton, and Oudin</label><mixed-citation>
      
Garcia, F., Folton, N., and Oudin, L.: Which objective function to calibrate
rainfall–runoff models for low-flow index simulations?,
Hydrolog. Sci. J., 62, 1149–1166,
<a href="https://doi.org/10.1080/02626667.2017.1308511" target="_blank">https://doi.org/10.1080/02626667.2017.1308511</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Gnann et al.(2021)Gnann, Coxon, Woods, Howden, and
McMillan</label><mixed-citation>
      
Gnann, S. J., Coxon, G., Woods, R. A., Howden, N. J., and McMillan, H. K.:
TOSSH: A Toolbox for Streamflow Signatures in Hydrology, Environ.
Modell. Softw., 138, 104983, <a href="https://doi.org/10.1016/j.envsoft.2021.104983" target="_blank">https://doi.org/10.1016/j.envsoft.2021.104983</a>,
2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Gupta et al.(2008)Gupta, Wagener, and
Liu</label><mixed-citation>
      
Gupta, H. V., Wagener, T., and Liu, Y.: Reconciling theory with observations:
elements of a diagnostic approach to model evaluation, Hydrol.
Process., 22, 3802–3813, <a href="https://doi.org/10.1002/hyp.6989" target="_blank">https://doi.org/10.1002/hyp.6989</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Gupta et al.(2009)Gupta, Kling, Yilmaz, and Martinez</label><mixed-citation>
      
Gupta, H. V., Kling, H., Yilmaz, K. K., and Martinez, G. F.: Decomposition of
the mean squared error and NSE performance criteria: Implications for
improving hydrological modelling, J. Hydrol., 377, 80–91,
<a href="https://doi.org/10.1016/j.jhydrol.2009.08.003" target="_blank">https://doi.org/10.1016/j.jhydrol.2009.08.003</a>, 2009.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Gupta et al.(2014)Gupta, Perrin, Blöschl, Montanari, Kumar, Clark,
and Andréassian</label><mixed-citation>
      
Gupta, H. V., Perrin, C., Blöschl, G., Montanari, A., Kumar, R., Clark, M., and Andréassian, V.: Large-sample hydrology: a need to balance depth with breadth, Hydrol. Earth Syst. Sci., 18, 463–477, <a href="https://doi.org/10.5194/hess-18-463-2014" target="_blank">https://doi.org/10.5194/hess-18-463-2014</a>, 2014.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Hallouin et al.(2020)</label><mixed-citation>
      
Hallouin, T., Bruen, M., and O'Loughlin, F. E.: Calibration of hydrological models for ecologically relevant streamflow predictions: a trade-off between fitting well to data and estimating consistent parameter sets?, Hydrol. Earth Syst. Sci., 24, 1031–1054, <a href="https://doi.org/10.5194/hess-24-1031-2020" target="_blank">https://doi.org/10.5194/hess-24-1031-2020</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Hansen et al.(2003)Hansen, Müller, and Koumoutsakos</label><mixed-citation>
      
Hansen, N., Müller, S. D., and Koumoutsakos, P.: Reducing the Time Complexity
of the Derandomized Evolution Strategy with Covariance Matrix Adaptation
(CMA-ES), Evol. Comput., 11, 1–18,
<a href="https://doi.org/10.1162/106365603321828970" target="_blank">https://doi.org/10.1162/106365603321828970</a>, 2003.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Hernandez-Suarez et al.(2018)Hernandez-Suarez, Nejadhashemi, Kropp,
Abouali, Zhang, and Deb</label><mixed-citation>
      
Hernandez-Suarez, J. S., Nejadhashemi, A. P., Kropp, I. M., Abouali, M., Zhang,
Z., and Deb, K.: Evaluation of the impacts of hydrologic model calibration
methods on predictability of ecologically-relevant hydrologic indices,
J. Hydrol., 564, 758–772, <a href="https://doi.org/10.1016/j.jhydrol.2018.07.056" target="_blank">https://doi.org/10.1016/j.jhydrol.2018.07.056</a>,
2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Jackson et al.(2019)Jackson, Roberts, Nelsen, Williams, Nelson, and
Ames</label><mixed-citation>
      
Jackson, E. K., Roberts, W., Nelsen, B., Williams, G. P., Nelson, E. J., and
Ames, D. P.: Introductory overview: Error metrics for hydrologic modelling
– A review of common practices and an open source library to facilitate use
and adoption, Environ. Modell. Softw., 119, 32–48,
<a href="https://doi.org/10.1016/j.envsoft.2019.05.001" target="_blank">https://doi.org/10.1016/j.envsoft.2019.05.001</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>Jacobs et al.(2024)Jacobs, Tobi, and Hengeveld</label><mixed-citation>
      
Jacobs, B., Tobi, H., and Hengeveld, G. M.: Linking error measures to model
questions, Ecol. Model., 487, 110562,
<a href="https://doi.org/10.1016/j.ecolmodel.2023.110562" target="_blank">https://doi.org/10.1016/j.ecolmodel.2023.110562</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Jakeman et al.(2006)Jakeman, Letcher, and Norton</label><mixed-citation>
      
Jakeman, A., Letcher, R., and Norton, J.: Ten iterative steps in development
and evaluation of environmental models, Environ. Modell. Softw.,
21, 602–614, <a href="https://doi.org/10.1016/j.envsoft.2006.01.004" target="_blank">https://doi.org/10.1016/j.envsoft.2006.01.004</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Jakeman et al.(2018)Jakeman, Sawah, Cuddy, Robson, McIntyre, and
Cook</label><mixed-citation>
      
Jakeman, A. J., Sawah, S. E., Cuddy, S., Robson, B., McIntyre, N., and Cook,
F.: Good Modelling Practice Principles, Tech. rep., The State of Queensland (Department of Environment and Science, <a href="https://science.desi.qld.gov.au/__data/assets/pdf_file/0011/81110/qwmn-good-modelling-practice-principles.pdf" target="_blank"/> (last access: 23 July 2026), 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Kiraz et al.(2023)Kiraz, Coxon, and Wagener</label><mixed-citation>
      
Kiraz, M., Coxon, G., and Wagener, T.: A Signature‐Based Hydrologic
Efficiency Metric for Model Calibration and Evaluation in Gauged and Ungauged
Catchments, Water Resour. Res., 59, <a href="https://doi.org/10.1029/2023WR035321" target="_blank">https://doi.org/10.1029/2023WR035321</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib34"><label>Kirchner(2006)</label><mixed-citation>
      
Kirchner, J. W.: Getting the right answers for the right reasons: Linking
measurements, analyses, and models to advance the science of hydrology, Water
Resour. Res., 42, <a href="https://doi.org/10.1029/2005WR004362" target="_blank">https://doi.org/10.1029/2005WR004362</a>, 2006.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib35"><label>Klingler et al.(2021)Klingler, Schulz, and
Herrnegger</label><mixed-citation>
      
Klingler, C., Schulz, K., and Herrnegger, M.: LamaH-CE: LArge-SaMple DAta for Hydrology and Environmental Sciences for Central Europe, Earth Syst. Sci. Data, 13, 4529–4565, <a href="https://doi.org/10.5194/essd-13-4529-2021" target="_blank">https://doi.org/10.5194/essd-13-4529-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib36"><label>Knoben(2024)</label><mixed-citation>
      
Knoben, W. J. M.: Setting expectations for hydrologic model performance with
an ensemble of simple benchmarks, Hydrol. Process., 38,
<a href="https://doi.org/10.1002/hyp.15288" target="_blank">https://doi.org/10.1002/hyp.15288</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib37"><label>Knoben et al.(2018)Knoben, Woods, and Freer</label><mixed-citation>
      
Knoben, W. J. M., Woods, R. A., and Freer, J. E.: A Quantitative Hydrological
Climate Classification Evaluated With Independent Streamflow Data, Water
Resour. Res., 54, 5088–5109, <a href="https://doi.org/10.1029/2018WR022913" target="_blank">https://doi.org/10.1029/2018WR022913</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib38"><label>Knoben et al.(2019a)Knoben, Freer, Fowler, Peel, and
Woods</label><mixed-citation>
      
Knoben, W. J. M., Freer, J. E., Fowler, K. J. A., Peel, M. C., and Woods, R. A.: Modular Assessment of Rainfall–Runoff Models Toolbox (MARRMoT) v1.2: an open-source, extendable framework providing implementations of 46 conceptual hydrologic models as continuous state-space formulations, Geosci. Model Dev., 12, 2463–2480, <a href="https://doi.org/10.5194/gmd-12-2463-2019" target="_blank">https://doi.org/10.5194/gmd-12-2463-2019</a>, 2019a.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib39"><label>Knoben et al.(2019b)Knoben, Freer, and
Woods</label><mixed-citation>
      
Knoben, W. J. M., Freer, J. E., and Woods, R. A.: Technical note: Inherent benchmark or not? Comparing Nash–Sutcliffe and Kling–Gupta efficiency scores, Hydrol. Earth Syst. Sci., 23, 4323–4331, <a href="https://doi.org/10.5194/hess-23-4323-2019" target="_blank">https://doi.org/10.5194/hess-23-4323-2019</a>, 2019b.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib40"><label>Knoben et al.(2020)Knoben, Freer, Peel, Fowler, and
Woods</label><mixed-citation>
      
Knoben, W. J. M., Freer, J. E., Peel, M. C., Fowler, K. J. A., and Woods,
R. A.: A Brief Analysis of Conceptual Model Structure Uncertainty Using 36
Models and 559 Catchments, Water Resour. Res., 56,
<a href="https://doi.org/10.1029/2019wr025975" target="_blank">https://doi.org/10.1029/2019wr025975</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib41"><label>Kratzert et al.(2023)Kratzert, Nearing, Addor, Erickson, Gauch,
Gilon, Gudmundsson, Hassidim, Klotz, Nevo, Shalev, and
Matias</label><mixed-citation>
      
Kratzert, F., Nearing, G., Addor, N., Erickson, T., Gauch, M., Gilon, O.,
Gudmundsson, L., Hassidim, A., Klotz, D., Nevo, S., Shalev, G., and Matias,
Y.: Caravan – A global community dataset for large-sample hydrology,
Scientific Data, 10, <a href="https://doi.org/10.1038/s41597-023-01975-w" target="_blank">https://doi.org/10.1038/s41597-023-01975-w</a>, 2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib42"><label>Krause et al.(2005)Krause, Boyle, and Bäse</label><mixed-citation>
      
Krause, P., Boyle, D. P., and Bäse, F.: Comparison of different efficiency criteria for hydrological model assessment, Adv. Geosci., 5, 89–97, <a href="https://doi.org/10.5194/adgeo-5-89-2005" target="_blank">https://doi.org/10.5194/adgeo-5-89-2005</a>, 2005.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib43"><label>Mai(2023)</label><mixed-citation>
      
Mai, J.: Ten strategies towards successful calibration of environmental models,
J. Hydrol., 620, 129414, <a href="https://doi.org/10.1016/j.jhydrol.2023.129414" target="_blank">https://doi.org/10.1016/j.jhydrol.2023.129414</a>,
2023.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib44"><label>McMillan(2020)</label><mixed-citation>
      
McMillan, H.: Linking hydrologic signatures to hydrologic processes: A review,
Hydrol. Process., 34, 1393–1409, <a href="https://doi.org/10.1002/hyp.13632" target="_blank">https://doi.org/10.1002/hyp.13632</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib45"><label>Melsen et al.(2019)Melsen, Teuling, Torfs, Zappa, Mizukami, Mendoza,
Clark, and Uijlenhoet</label><mixed-citation>
      
Melsen, L. A., Teuling, A. J., Torfs, P. J., Zappa, M., Mizukami, N., Mendoza,
P. A., Clark, M. P., and Uijlenhoet, R.: Subjective modeling decisions can
significantly impact the simulation of flood and drought events, J.
Hydrol., 568, 1093–1104, <a href="https://doi.org/10.1016/j.jhydrol.2018.11.046" target="_blank">https://doi.org/10.1016/j.jhydrol.2018.11.046</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib46"><label>Melsen et al.(2025)Melsen, Puy, Torfs, and
Saltelli</label><mixed-citation>
      
Melsen, L. A., Puy, A., Torfs, P. J. J. F., and Saltelli, A.: The rise of the
Nash-Sutcliffe efficiency in hydrology, Hydrolog. Sci. J.,
70, 1248–1259, <a href="https://doi.org/10.1080/02626667.2025.2475105" target="_blank">https://doi.org/10.1080/02626667.2025.2475105</a>, 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib47"><label>Mendoza et al.(2016)Mendoza, Clark, Mizukami, Gutmann, Arnold,
Brekke, and Rajagopalan</label><mixed-citation>
      
Mendoza, P. A., Clark, M. P., Mizukami, N., Gutmann, E. D., Arnold, J. R.,
Brekke, L. D., and Rajagopalan, B.: How do hydrologic modeling decisions
affect the portrayal of climate change impacts?, Hydrol. Process., 30,
1071–1095, <a href="https://doi.org/10.1002/hyp.10684" target="_blank">https://doi.org/10.1002/hyp.10684</a>, 2016.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib48"><label>Mizukami et al.(2019)Mizukami, Rakovec, Newman, Clark, Wood, Gupta,
and Kumar</label><mixed-citation>
      
Mizukami, N., Rakovec, O., Newman, A. J., Clark, M. P., Wood, A. W., Gupta, H. V., and Kumar, R.: On the choice of calibration metrics for “high-flow” estimation using hydrologic models, Hydrol. Earth Syst. Sci., 23, 2601–2614, <a href="https://doi.org/10.5194/hess-23-2601-2019" target="_blank">https://doi.org/10.5194/hess-23-2601-2019</a>, 2019.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib49"><label>Nash and Sutcliffe(1970)</label><mixed-citation>
      
Nash, J. and Sutcliffe, J.: River flow forecasting through conceptual models
part I – A discussion of principles, J. Hydrol., 10,
282–290, <a href="https://doi.org/10.1016/0022-1694(70)90255-6" target="_blank">https://doi.org/10.1016/0022-1694(70)90255-6</a>, 1970.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib50"><label>Nemri and Kinnard(2020)</label><mixed-citation>
      
Nemri, S. and Kinnard, C.: Comparing calibration strategies of a conceptual
snow hydrology model and their impact on model performance and parameter
identifiability, J. Hydrol., 582, 124474,
<a href="https://doi.org/10.1016/j.jhydrol.2019.124474" target="_blank">https://doi.org/10.1016/j.jhydrol.2019.124474</a>, 2020.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib51"><label>Pool et al.(2017)Pool, Vis, Knight, and Seibert</label><mixed-citation>
      
Pool, S., Vis, M. J. P., Knight, R. R., and Seibert, J.: Streamflow characteristics from modeled runoff time series – importance of calibration criteria selection, Hydrol. Earth Syst. Sci., 21, 5443–5457, <a href="https://doi.org/10.5194/hess-21-5443-2017" target="_blank">https://doi.org/10.5194/hess-21-5443-2017</a>, 2017.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib52"><label>Pool et al.(2018)Pool, Vis, and
Seibert</label><mixed-citation>
      
Pool, S., Vis, M., and Seibert, J.: Evaluating model performance: towards a
non-parametric variant of the Kling-Gupta efficiency, Hydrolog.
Sci. J., 63, 1941–1953, <a href="https://doi.org/10.1080/02626667.2018.1552002" target="_blank">https://doi.org/10.1080/02626667.2018.1552002</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib53"><label>Santos et al.(2018)Santos, Thirel, and
Perrin</label><mixed-citation>
      
Santos, L., Thirel, G., and Perrin, C.: Technical note: Pitfalls in using log-transformed flows within the KGE criterion, Hydrol. Earth Syst. Sci., 22, 4583–4591, <a href="https://doi.org/10.5194/hess-22-4583-2018" target="_blank">https://doi.org/10.5194/hess-22-4583-2018</a>, 2018.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib54"><label>Schaefli and Gupta(2007)</label><mixed-citation>
      
Schaefli, B. and Gupta, H. V.: Do Nash values have value?, Hydrol.
Process., 21, 2075–2080, <a href="https://doi.org/10.1002/hyp.6825" target="_blank">https://doi.org/10.1002/hyp.6825</a>, 2007.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib55"><label>Schwemmle et al.(2021)Schwemmle, Demand, and Weiler</label><mixed-citation>
      
Schwemmle, R., Demand, D., and Weiler, M.: Technical note: Diagnostic efficiency – specific evaluation of model performance, Hydrol. Earth Syst. Sci., 25, 2187–2198, <a href="https://doi.org/10.5194/hess-25-2187-2021" target="_blank">https://doi.org/10.5194/hess-25-2187-2021</a>, 2021.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib56"><label>Seibert(2001)</label><mixed-citation>
      
Seibert, J.: On the need for benchmarks in hydrological modelling, Hydrol.
Process., 15, 1063–1064, <a href="https://doi.org/10.1002/hyp.446" target="_blank">https://doi.org/10.1002/hyp.446</a>, 2001.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib57"><label>Seiller et al.(2017)Seiller, Roy, and Anctil</label><mixed-citation>
      
Seiller, G., Roy, R., and Anctil, F.: Influence of three common calibration
metrics on the diagnosis of climate change impacts on water resources,
J. Hydrol., 547, 280–295, <a href="https://doi.org/10.1016/j.jhydrol.2017.02.004" target="_blank">https://doi.org/10.1016/j.jhydrol.2017.02.004</a>,
2017.


    </mixed-citation></ref-html>
<ref-html id="bib1.bib58"><label>Thirel et al.(2024)Thirel, Santos, Delaigue, and Perrin</label><mixed-citation>
      
Thirel, G., Santos, L., Delaigue, O., and Perrin, C.: On the use of streamflow transformations for hydrological model calibration, Hydrol. Earth Syst. Sci., 28, 4837–4860, <a href="https://doi.org/10.5194/hess-28-4837-2024" target="_blank">https://doi.org/10.5194/hess-28-4837-2024</a>, 2024.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib59"><label>Trotter and Knoben(2022)</label><mixed-citation>
      
Trotter, L. and Knoben, W. J. M.: MARRMoT v2.1, Zenodo [code], <a href="https://doi.org/10.5281/zenodo.6484372" target="_blank">https://doi.org/10.5281/zenodo.6484372</a>,
2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib60"><label>Trotter et al.(2022)Trotter, Knoben, Fowler, Saft, and
Peel</label><mixed-citation>
      
Trotter, L., Knoben, W. J. M., Fowler, K. J. A., Saft, M., and Peel, M. C.: Modular Assessment of Rainfall–Runoff Models Toolbox (MARRMoT) v2.1: an object-oriented implementation of 47 established hydrological models for improved speed and readability, Geosci. Model Dev., 15, 6359–6369, <a href="https://doi.org/10.5194/gmd-15-6359-2022" target="_blank">https://doi.org/10.5194/gmd-15-6359-2022</a>, 2022.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib61"><label>van Waveren et al.(1999)Waveren, Groot, Scholten, Geer, Wösten, Koeze,
and Noort</label><mixed-citation>
      
van Waveren, R. H., Groot, S., Scholten, H., Geer, F. v., Wösten, J., Koeze, R.,
and Noort, J.: Good Modelling Practice Handbook, Tech. rep., Dutch Dept. of Public Works, Institute for Inland Water Management and Waste Water Treatment, report 99.036, ISBN 9057730561, 1999.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib62"><label>van Werkhoven et al.(2008)van Werkhoven, Wagener, Reed, and
Tang</label><mixed-citation>
      
van Werkhoven, K., Wagener, T., Reed, P., and Tang, Y.: Characterization of watershed model behavior across a hydroclimatic gradient, Water Resour. Res., 44, <a href="https://doi.org/10.1029/2007WR006271" target="_blank">https://doi.org/10.1029/2007WR006271</a>, 2008.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib63"><label>Vis et al.(2015)Vis, Knight, Pool, Wolfe, and Seibert</label><mixed-citation>
      
Vis, M., Knight, R., Pool, S., Wolfe, W., and Seibert, J.: Model Calibration
Criteria for Estimating Ecological Flow Characteristics, Water, 7,
2358–2381, <a href="https://doi.org/10.3390/w7052358" target="_blank">https://doi.org/10.3390/w7052358</a>, 2015.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib64"><label>Wagener and
Montanari(2011)</label><mixed-citation>
      
Wagener, T. and Montanari, A.: Convergence of approaches toward reducing
uncertainty in predictions in ungauged basins, Water Resour. Res., 47,
2010WR009469, <a href="https://doi.org/10.1029/2010WR009469" target="_blank">https://doi.org/10.1029/2010WR009469</a>, 2011.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib65"><label>Wagener(2025)</label><mixed-citation>
      
Wagener, P.: Code for Analysis and Visualization, GitHub [code],
<a href="https://github.com/peterwagener/OF_Signature_Code.git" target="_blank"/> (last access: 16 July 2026), 2025.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib66"><label>Wagener(2026)</label><mixed-citation>
      
Wagener, P.: Calibration and Evaluation Data for Objective Function Impact on Hydrologic Signatures, Zenodo [data set],
<a href="https://doi.org/10.5281/zenodo.21399824" target="_blank">https://doi.org/10.5281/zenodo.21399824</a>, 2026.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib67"><label>Yapo et al.(1998)Yapo, Gupta, and
Sorooshian</label><mixed-citation>
      
Yapo, P. O., Gupta, H. V., and Sorooshian, S.: Multi-objective global
optimization for hydrologic models, J. Hydrol., 204, 83–97,
<a href="https://doi.org/10.1016/S0022-1694(97)00107-8" target="_blank">https://doi.org/10.1016/S0022-1694(97)00107-8</a>, 1998.

    </mixed-citation></ref-html>
<ref-html id="bib1.bib68"><label>Yilmaz et al.(2008)Yilmaz, Gupta, and
Wagener</label><mixed-citation>
      
Yilmaz, K. K., Gupta, H. V., and Wagener, T.: A process‐based diagnostic
approach to model evaluation: Application to the NWS distributed
hydrologic model, Water Resour. Res., 44, 2007WR006716,
<a href="https://doi.org/10.1029/2007WR006716" target="_blank">https://doi.org/10.1029/2007WR006716</a>, 2008.

    </mixed-citation></ref-html>--></article>
