Articles | Volume 30, issue 14
https://doi.org/10.5194/hess-30-4629-2026
© Author(s) 2026. This work is distributed under the Creative Commons Attribution 4.0 License.
Benchmarking reservoir operation schemes for large-scale hydrological models
Download
- Final revised paper (published on 22 Jul 2026)
- Preprint (discussion started on 24 Feb 2026)
Interactive discussion
Status: closed
Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor
| : Report abuse
-
RC1: 'Comment on egusphere-2026-904', Saskia Salwey, 23 Mar 2026
- AC1: 'Reply on RC1', Jesús Casado Rodríguez, 28 Apr 2026
-
RC2: 'Comment on egusphere-2026-904', Anonymous Referee #2, 30 Mar 2026
- AC2: 'Reply on RC2', Jesús Casado Rodríguez, 28 Apr 2026
Peer review completion
AR – Author's response | RR – Referee report | ED – Editor decision | EF – Editorial file upload
ED: Publish subject to revisions (further review by editor and referees) (09 May 2026) by Fanny Sarrazin
AR by Jesús Casado Rodríguez on behalf of the Authors (15 Jun 2026)
Author's response
Author's tracked changes
Manuscript
ED: Referee Nomination & Report Request started (20 Jun 2026) by Fanny Sarrazin
ED: Publish subject to minor revisions (review by editor) (12 Jul 2026) by Fanny Sarrazin
AR by Jesús Casado Rodríguez on behalf of the Authors (13 Jul 2026)
Author's response
Author's tracked changes
Manuscript
ED: Publish as is (13 Jul 2026) by Fanny Sarrazin
AR by Jesús Casado Rodríguez on behalf of the Authors (16 Jul 2026)
Manuscript
This study benchmarks the performance of four reservoir operation schemes across the US in a large-scale hydrological model. The manuscript compares four different calibration strategies, finding that calibrating to reservoir storage is more informative than calibrating to reservoir outflow. I think that the results of this study are very important for the modelling community and the take-home messages should be carefully considered by anyone incorporating reservoirs into large-scale hydrological modelling. In fact, I have been hoping someone would publish a study similar to this for a while so thank you! In general the manuscript was very clear and well-written, but below I have left a few suggestions for how I think it could be improved.
It would be interesting to hear more about your model calibration strategy (as introduced on L60). It is not super clear to me how the specific details of the calibration worked. Am I right in thinking that you used the ‘default’ reservoir parameters from the literature listed in Tables 2, 3 and 4 and then calibrated the non-reservoir parameters around these? If so, considering that the results using the default reservoir parameters often failed to capture the storage dynamics well, do you think that the non-reservoir parameters were calibrated in a way which means they overcompensate for poorly represented reservoir processes? Did you compare the selected non-reservoir parameters to values used in natural catchments or the literature to see whether they were physically realistic? Perhaps you can elaborate on this a bit.
Finding that calibrating the model to reservoir storage is more informative than calibrating to outflow is a really interesting (and useful!) result. I am pleased that your results suggest we may be able to utilize satellite data for model calibration but wonder whether you should demonstrate this in the manuscript. Did you try integrating satellite data (e.g. the data you discuss in Appendix A) into some of your reservoir storage calibration experiments? If not I think this would be a very valuable addition to Appendix A or the manuscript. If this paper is going to advocate for this possibility it would be nice to showcase this, particularly because in many places storage data like in ResOpsUS is not available. It would be interesting to know how the differences in satellite derived storage impact the results.
How were the thresholds for DOR and DOD selected to define a significantly altered natural flow regime?
At some point (even if in an appendix) I would be interested to hear about the breakdown of the individual KGE components. Did all aspects of the metric perform similarly or were there some that were always high or always low? How did this vary across reservoirs of different types?
Could Figure1 also show the primary purpose of the reservoirs? This seems important for the operations. Is it possible to add some analysis somewhere which describes how the KGE performance varied across reservoirs of different types? This could link nicely to the discussion in section 5.2. I think it would be useful for readers to understand whether the results of this study would apply to other locations where perhaps there is a different distribution of reservoir types.
I think one of the most interesting results in this paper is on L356 where you state that STARFIT was not markedly superior to CaMa-Flood which is far simpler. There is often an assumption in our field that more data/ complexity will always lead to better results and so I think it is important that we highlight that this is not always the case. Could you consider mentioning this in the abstract?
You mention several times that STARFIT still has distinct advantages over the other schemes (e.g. on L357 and L445) but I cannot see how your results evidence this? It seems to have been outperformed by simpler methods. Can you make it clearer why you think this?
I think it is really nice that you have published the ResOpsUS+CARS dataset!