Articles | Volume 30, issue 14
https://doi.org/10.5194/hess-30-4667-2026
© Author(s) 2026. This work is distributed under the Creative Commons Attribution 4.0 License.
Hybrid models generalize better to warmer climate conditions than process-based and purely data-driven models
Download
- Final revised paper (published on 27 Jul 2026)
- Preprint (discussion started on 07 Nov 2025)
Interactive discussion
Status: closed
Comment types: AC – author | RC – referee | CC – community | EC – editor | CEC – chief editor
| : Report abuse
-
RC1: 'Comment on egusphere-2025-5201', Anonymous Referee #1, 04 Dec 2025
- AC1: 'Reply on RC1', Raul R. Wood, 19 Feb 2026
-
CC1: 'How do we define 'generalizability'?', Sacha Ruzzante, 17 Dec 2025
- AC3: 'Reply on CC1', Raul R. Wood, 19 Feb 2026
-
RC2: 'Comment on egusphere-2025-5201', Anonymous Referee #2, 20 Jan 2026
- AC2: 'Reply on RC2', Raul R. Wood, 19 Feb 2026
Peer review completion
AR – Author's response | RR – Referee report | ED – Editor decision | EF – Editorial file upload
ED: Publish subject to revisions (further review by editor and referees) (07 Mar 2026) by Elena Toth
AR by Raul R. Wood on behalf of the Authors (30 Apr 2026)
Author's response
Author's tracked changes
Manuscript
ED: Referee Nomination & Report Request started (30 Apr 2026) by Elena Toth
RR by Anonymous Referee #2 (09 Jun 2026)
ED: Publish as is (11 Jun 2026) by Elena Toth
AR by Raul R. Wood on behalf of the Authors (19 Jun 2026)
The manuscript evaluates the ability of different model types (HBV, LSTM, and a hybrid model) to predict river streamflow under different climate conditions, particularly when the training/calibration period differs from the testing/validation/prediction period. This issue is critical when applying machine learning models to future climate-change impact studies. The manuscript is well written, the experimental design is appropriate for the scientific questions, and the results are clearly illustrated. I have several major comments that I would like to discuss with the authors. If these can be addressed, I would recommend the paper for publication.
First, I think a sensitivity test should be conducted. Before applying the models to the warmer period, perturb the input variables (such as temperature or precipitation) and evaluate how the models respond to these changes. This is relevent for the following analysis, maybe different model is sensitive, others are not.
Another concern relates to the importance of the different input features. Is temperature the most important predictor, or do other variables differ more between the cold and warm periods? The manuscript does not discuss precipitation changes, and I think a feature-importance/SHAP analysis is possible for the LSTM or hybrid model. It would be helpful to understand whether precipitation or PET, although changing less than temperature, may have a stronger influence on streamflow. Concerning the evaluation metrics are not very different from models to models during different period. just to confirm that different model performances are due to climate warming.
A few minor comments
Line 1: Use consistent terminology: either “deep learning,” “deep-learning” (as an adjective), or “DL,” throughout the manuscript.
Line 1: Spell out “Long Short-Term Memory (LSTM)” on first use.
The abstract is currently very conceptual. Please include key numerical results (e.g., NSE, KGE) to quantify performance. For example, Lines 10–12 mention that the LSTM performs best during the cold period but worse during the warm period, this should be supported with specific numbers.
From the abstract, the advantages of the hybrid model over the LSTM are not obvious. Lines 10–15 suggest that hybrid models have similar accuracy to LSTMs, please clarify the added benefit.
Line 86: Please correct the citation formatting.
The introduction is well written.
Line 153: If the Po River basin is not included in the analysis, it may be better not to mention it here (or clarify this later, as in Line 159).
Line 310: Please clarify the distinction between “in-sample HBV” and “regional HBV.”
Line 320: This relates to my major concern, how does precipitation change between periods and among different catchments?
Lines 315–319: The reported values are very close to each other, and they represent means or medians over hundreds of catchments. Could these differences fall within model uncertainty?
Section 3.4: In general, the hybrid and HBV models perform worse than the LSTM model. Is this due to limitations of HBV in snow-affected catchments, where LSTM may better learn snow–streamflow relationships? Does the hybrid model inherit these limitations from HBV, preventing it from outperforming the LSTM?
Line 415: All models show higher performance for flood events than for drought or low flows. Is this due to the choice of objective function (NSE), which emphasizes high-flow periods?