Articles | Volume 30, issue 18
https://doi.org/10.5194/hess-30-5857-2026
© Author(s) 2026. This work is distributed under the Creative Commons Attribution 4.0 License.
Setting the bar: benchmarks for model performances in large-sample hydrology
Download
- Final revised paper (published on 16 Sep 2026)
- Supplement to the final revised paper
- Preprint (discussion started on 11 Jun 2026)
- Supplement to the preprint
Interactive discussion
Status: closed
Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor
| : Report abuse
-
RC1: 'Very valuable, but needs to provide details on setup and performance values', Benedikt Heudorfer, 19 Jun 2026
- AC1: 'Quick Reply on link comment in RC1', Jan Seibert, 22 Jun 2026
- AC2: 'Reply on RC1', Jan Seibert, 08 Jul 2026
-
RC2: 'Comment on egusphere-2026-3272', Tam Nguyen, 24 Jun 2026
- AC3: 'Reply on RC2', Jan Seibert, 08 Jul 2026
Peer review completion
AR – Author's response | RR – Referee report | ED – Editor decision | EF – Editorial file upload
ED: Publish subject to technical corrections (30 Jul 2026) by Ralf Loritz
AR by Jan Seibert on behalf of the Authors (05 Sep 2026)
Author's response
Manuscript
General comment:
This paper deals with very relevant questions w.r.t. proper comparison of model quality on large sample hydrological datasets. It uses the HBV conceptual model to provide upper and lower performance boundaries on dataset and catchment basis. This is a highly relevant contribution that can also guide the general machine learning community as it provides (among other things) reasonable lower boundaries of predictability of individual catchments from a conceptual hydrological model. Since neural networks are inherently random, the performance of the model can vary substantially between for each catchment between model runs. The study presented here can aid this shortcoming by telling us where good performance can be expected and where not.
Major comment:
However the main shortcoming of this paper is that it needs to provide better documentation of how the models are actually run, on which catchments they are run, on which time periods they are trained, what the exact performance is for each respective dataset and/or catchment, etc.. While this is (mostly) not sensible to report in the main body of the paper, these details should be included in the supplement files, in a data repository, published along with the code, anywhere (the link in the data availability section is dead). These details are crucial to allow what the study sets out to do: enabling comparison/benchmarking of models in future studies, because without these details, future studies will not be able to replicate the experimental setup properly. See also specific comments on this topic.
Specific comments:
Line 46 introduce NPE properly please.
Line 98: dot missing
Section 2.3 is missing crucial information e.g. on test periods. To allow future studies to compare performance, the exact test periods should be named for each dataset. Backgroung: The choice of data is really important how the value of the performance metric turns out (as the “elephant in the room” paper my Maier 2023 points out: https://doi.org/10.1016/j.envsoft.2023.105779). If exact periods and catchment lists are not reported for each dataset, valid comparison with your results are practically impossible, rendering the attempt to provide benchmark for future reference futile. Also, exact catchment list should be reported somewhere, on which you train. This is needed for people to match their model setup exactly in the future, in case they want to benchmark the results.
Line 117-119: “The first one to two years were used as a warming-up period to obtain reasonable initial conditions for the storage components of HBV. The remaining years were then used for model calibration (upper benchmark) and evaluation (lower benchmark) with streamflow.” Does that mean no test period for calibrated models, and the reported upper benchmarks are from the calibration period metrics? If yes, that should be stated clearly. I don’t know how it is treated in the conceptual modelling community, but in machine learning this would be considered data leakage and not a reliable metric for model performance, as it will mostly speak for the degree of overfitting. Could you please elaborate whether the phenomenon of overfitting is relevant in the conceptual modelling domain, or why the choice to report calibration period metric makes sense? Or if I misunderstand the section and this all is beside the point, please clarify.
Line 131-134: very good strategy and reflects the authors knowledge about the dependency of metrics to calculation procedure. Speaks for the rigidity of the study.
135-138: here as well, the exact catchment IDs associated with the lower benchmark should be made public somewhere, if not already present in a published codebase. That is, if the random selection of the 10 catchments was not repeated with replacement, in which case tracing it may be pointless. Backgroung: The choice of data is really important how the value of the performance metric turns out.
Section 3.1: only evaluation based on NPE is reported, but 2.3.1 claims it was evaluated based on NSE, KGE and NPE; since NPE is more novel and NSE and KGE are more widely used, KGE and NSe should be reported as well; also, to allow proper use of this benchmark in future studies, tables with exact metric values (incl. any uncertainty of distribution indices like IQR etc.) should be reported, if only in the supplements.
Section 3.2: Having this effect of the influence of the calculation procedure on the final metric value finally worked out in a paper is super exiting to me. We have briefly discussed this effect in our 2025 paper (https://doi.org/10.1029/2024GL113036) and internal tests show metric value deviations of up to 0.1 NSE depending on the calculation procedure, but I am not aware that this was specifically addressed anywhere yet. Very nice result, can be highlighted more prominently in my humble opinion.
Line 208-209: I don’t understand exactly how this parameter range width is applied; how can a negative parameter range be? Also range values should be included in the figure 3 xlabs as well.
Table 3: ah, here is the table on exact upper and lower benchmark values ; but should be reported further split into the individual datasets, so people can compare their models against it in the future. These are surprisingly strong lower benchmark by the way. Reading this paper made me think how a lower benchmark might be computed for ML/DL models. If you have any ideas, I’d appreciate elaborating and/or putting them into the paper.