<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing with OASIS Tables v3.0 20080202//EN" "journalpub-oasis3.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:oasis="http://docs.oasis-open.org/ns/oasis-exchange/table" xml:lang="en" dtd-version="3.0">
  <front>
    <journal-meta><journal-id journal-id-type="publisher">HESS</journal-id><journal-title-group>
    <journal-title>Hydrology and Earth System Sciences</journal-title>
    <abbrev-journal-title abbrev-type="publisher">HESS</abbrev-journal-title><abbrev-journal-title abbrev-type="nlm-ta">Hydrol. Earth Syst. Sci.</abbrev-journal-title>
  </journal-title-group><issn pub-type="epub">1607-7938</issn><publisher>
    <publisher-name>Copernicus Publications</publisher-name>
    <publisher-loc>Göttingen, Germany</publisher-loc>
  </publisher></journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.5194/hess-22-4867-2018</article-id><title-group><article-title>Cross-validation of bias-corrected climate simulations is misleading</article-title><alt-title>Misleading cross-validation of bias correction</alt-title>
      </title-group><?xmltex \runningtitle{Misleading cross-validation of bias correction}?><?xmltex \runningauthor{D.~Maraun and M.~Widmann}?>
      <contrib-group>
        <contrib contrib-type="author" corresp="yes" rid="aff1">
          <name><surname>Maraun</surname><given-names>Douglas</given-names></name>
          <email>douglas.maraun@uni-graz.at</email>
        </contrib>
        <contrib contrib-type="author" corresp="no" rid="aff2">
          <name><surname>Widmann</surname><given-names>Martin</given-names></name>
          
        <ext-link>https://orcid.org/0000-0001-5447-5763</ext-link></contrib>
        <aff id="aff1"><label>1</label><institution>Wegener Center for Climate and Global Change, University of Graz, Brandhofgasse 5, 8010 Graz, Austria</institution>
        </aff>
        <aff id="aff2"><label>2</label><institution>School of Geography, Earth and Environmental Sciences, University of Birmingham, Birmingham, B15 2TT, UK</institution>
        </aff>
      </contrib-group>
      <author-notes><corresp id="corr1">Douglas Maraun (douglas.maraun@uni-graz.at)</corresp></author-notes><pub-date><day>18</day><month>September</month><year>2018</year></pub-date>
      
      <volume>22</volume>
      <issue>9</issue>
      <fpage>4867</fpage><lpage>4873</lpage>
      <history>
        <date date-type="received"><day>26</day><month>March</month><year>2018</year></date>
           <date date-type="rev-request"><day>17</day><month>April</month><year>2018</year></date>
           <date date-type="rev-recd"><day>28</day><month>July</month><year>2018</year></date>
           <date date-type="accepted"><day>27</day><month>August</month><year>2018</year></date>
      </history>
      <permissions>
        
        
      <license license-type="open-access"><license-p>This work is licensed under the Creative Commons Attribution 4.0 International License. To view a copy of this licence, visit <ext-link ext-link-type="uri" xlink:href="https://creativecommons.org/licenses/by/4.0/">https://creativecommons.org/licenses/by/4.0/</ext-link></license-p></license></permissions><self-uri xlink:href="https://hess.copernicus.org/articles/22/4867/2018/hess-22-4867-2018.html">This article is available from https://hess.copernicus.org/articles/22/4867/2018/hess-22-4867-2018.html</self-uri><self-uri xlink:href="https://hess.copernicus.org/articles/22/4867/2018/hess-22-4867-2018.pdf">The full text article is available as a PDF file from https://hess.copernicus.org/articles/22/4867/2018/hess-22-4867-2018.pdf</self-uri>
      <abstract>
    <p id="d1e95">We demonstrate both analytically and with a modelling example that
cross-validation of free-running bias-corrected climate change simulations
against observations is misleading. The underlying reasoning is as follows: a
cross-validation can have in principle two outcomes. A negative (in the sense
of not rejecting a null hypothesis), if the residual bias in the validation
period after bias correction vanishes; and a positive, if the residual bias
in the validation period after bias correction is large. It can be shown
analytically that the residual bias depends solely on the difference between
the simulated and observed change between calibration and validation periods.
This change, however, depends mainly on the realizations of internal
variability in the observations and climate model. As a consequence, the
outcome of a cross-validation is also dominated by internal variability, and
does not allow for any conclusion about the sensibility of a bias correction.
In particular, a sensible bias correction may be rejected (false positive)
and a non-sensible bias correction may be accepted (false negative). We
therefore propose to avoid cross-validation when evaluating bias correction
of free-running bias-corrected climate change simulations against
observations. Instead, one should evaluate non-calibrated temporal, spatial
and process-based aspects.</p>
  </abstract>
    </article-meta>
  </front>
<body>
      

<sec id="Ch1.S1" sec-type="intro">
  <title>Introduction</title>
      <p id="d1e105">Bias correction is a widely used approach to postprocess climate model
simulations before they are applied in impact studies
<xref ref-type="bibr" rid="bib1.bibx5 bib1.bibx8 bib1.bibx6 bib1.bibx31" id="paren.1"><named-content content-type="pre">e.g.</named-content></xref>. A wide
range of different correction methods has been developed, ranging from simple
additive or multiplicative corrections to quantile-based approaches. For
reviews of bias correction see <xref ref-type="bibr" rid="bib1.bibx28" id="normal.2"/>, <xref ref-type="bibr" rid="bib1.bibx17" id="normal.3"/> and the
book by <xref ref-type="bibr" rid="bib1.bibx19" id="normal.4"/>.</p>
      <p id="d1e122">The performance of a bias correction is typically evaluated against
independent observational data, which have not entered the calibration of the
correction function. For instance, <xref ref-type="bibr" rid="bib1.bibx24 bib1.bibx25" id="normal.5"/>, <xref ref-type="bibr" rid="bib1.bibx13" id="normal.6"/> and
<xref ref-type="bibr" rid="bib1.bibx3" id="normal.7"/> apply the holdout method, i.e. they calibrate the method on a
calibration period and evaluate it on a non-overlapping validation period.
Some authors even apply a full cross-validation, most often by permuting
calibration and validation period <xref ref-type="bibr" rid="bib1.bibx7" id="paren.8"><named-content content-type="pre">a 2-fold
cross-validation;</named-content></xref>.</p>
      <p id="d1e139">Cross-validation is a well-known and widely used statistical concept to
assess the skill of predictive statistical models <xref ref-type="bibr" rid="bib1.bibx27 bib1.bibx4" id="paren.9"/>. It
has been successfully applied in the atmospheric sciences, e.g. in weather
forecasting <xref ref-type="bibr" rid="bib1.bibx11 bib1.bibx22 bib1.bibx32" id="paren.10"/> and perfect predictor
experiments of downscaling methods <xref ref-type="bibr" rid="bib1.bibx29 bib1.bibx20" id="paren.11"/>.</p>
      <p id="d1e151">In climate change applications, however, the setting is typically different
from a weather forecasting or perfect predictor setting: here, the model is
running free, i.e. only external forcings are common to observation and
simulation. Internal climate variability on all scales is independent and not
synchronized. In this setting,
the aim is not to assess predictive power, e.g. on a day-by-day or
season-by-season basis, as in weather forecasting – in fact, by construction
it cannot be assessed. Importantly, observed and simulated long-term trends
may also differ substantially, just because of different random realizations
of long-term modes of variability. Prominent examples of such modes are the
Pacific Decadal Oscillation <xref ref-type="bibr" rid="bib1.bibx14" id="paren.12"><named-content content-type="pre">PDO,</named-content></xref> and the Atlantic
Multidecadal Oscillation <xref ref-type="bibr" rid="bib1.bibx26" id="paren.13"><named-content content-type="pre">AMO,</named-content></xref>.</p>
      <?pagebreak page4868?><p id="d1e165">These differences have crucial implications for the application of
cross-validation or any evaluation on
independent data. Our results build upon a recent study by <xref ref-type="bibr" rid="bib1.bibx21" id="normal.14"/>,
who demonstrated that in a climate change setting cross-validation of
marginal aspects is not able to identify bias correction skill. Here, we
additionally show that the outcome of a cross-validation is essentially
random and independent of the sensibility of the bias correction. We will
demonstrate these consequences for the holdout method, but they can of course
be generalized to any type of cross-validation.</p>
      <p id="d1e171">We will discuss the specific context of climate change simulations in
Sect. <xref ref-type="sec" rid="Ch1.S2"/>. An analytical derivation of the
cross-validation problem will be given in Sect. <xref ref-type="sec" rid="Ch1.S3"/>,
a modelling example in Sect. <xref ref-type="sec" rid="Ch1.S4"/>. We will close with
a discussion of the implications of our findings.</p>
</sec>
<sec id="Ch1.S2">
  <title>Cross-validation in the climate context</title>
      <p id="d1e186">Cross-validation was developed to quantify the predictive skill of
statistical models in the 1930s, and has become widely used with the advent
of modern computers <xref ref-type="bibr" rid="bib1.bibx27 bib1.bibx4" id="paren.15"/>. It has become a standard tool in
weather and climate forecasting
<xref ref-type="bibr" rid="bib1.bibx23 bib1.bibx11 bib1.bibx32 bib1.bibx22" id="paren.16"/>.</p>
      <p id="d1e195">The first major aim of cross-validation is to eliminate artificial skill: if
the statistical model is evaluated on the same data that are used for
calibration, the performance to predict new data will almost certainly be
lower than the estimated skill. Hence, the model is calibrated only on a
subset of the data, and evaluated on another – ideally independent – subset
of the data. This so-called holdout method, however, uses each data point
only either for calibration or validation and thus suffers from relatively
high sampling errors.</p>
      <p id="d1e198">The second major aim of cross-validation is therefore to use the data
optimally. To this end, the holdout method, i.e. training and validation, is
repeated on different subsets of the data. The simplest approach is the
so-called split sample method, where the data are just split once into two subsets. More advanced
<inline-formula><mml:math id="M1" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>-fold cross-validation splits the data set
into <inline-formula><mml:math id="M2" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula> non-overlapping blocks; in each fold, <inline-formula><mml:math id="M3" display="inline"><mml:mi>k</mml:mi></mml:math></inline-formula>–<inline-formula><mml:math id="M4" display="inline"><mml:mn mathvariant="normal">1</mml:mn></mml:math></inline-formula> blocks are used for
calibration and the remaining block is used for validation.</p>
      <p id="d1e229">In weather and climate predictions, the aim is to predict the weather,
i.e. internal variability, with a given lead time (say, 3 days or a season)
at a desired timescale (say, 6 h or a season). A typical evaluation assesses
how well certain meteorological aspects are predicted: in weather
forecasting, one may for instance be interested in the overall prediction
accuracy, measured by the root-mean squared error between predicted and
observed daily time series. In a seasonal prediction, one may be interested
in the bias of the predicted mean, or in the bias of the predicted wet-day
frequency over a season. In this context, a cross-validation makes perfect
sense if the validation blocks are long compared to the prediction lead time
(and process memory).</p>
      <p id="d1e233">Downscaling and bias correction methods are typically tested in perfect
predictor or perfect boundary condition experiments <xref ref-type="bibr" rid="bib1.bibx20" id="paren.17"/>, where
predictors or boundary conditions are taken from reanalysis data. The aim of
the downscaling in this context is not to predict internal variability ahead
into the future, but rather to predict the local weather conditional on the
state of the large-scale weather (i.e. to simulate the correct local
long-term weather statistics). Still, in such a setting cross-validation
makes perfect sense: the choice of reanalysis data as predictors/boundary
conditions synchronizes simulated and observed local variability on
timescales beyond a few weeks, such that the evaluation framework is similar
to the case of seasonal prediction.</p>
      <p id="d1e239">In free-running climate simulations, however, the situation is fundamentally
different: here, any predictive power results only from external (e.g.
anthropogenic) forcing at very long timescales, but internal variability is
not synchronized at any timescale. Yet long-term modes of internal climate
variability, such as the PDO <xref ref-type="bibr" rid="bib1.bibx26" id="paren.18"/> and the AMO
<xref ref-type="bibr" rid="bib1.bibx26" id="paren.19"/>, often mask forced climate trends even at multidecadal
timescales <xref ref-type="bibr" rid="bib1.bibx2 bib1.bibx16" id="paren.20"/>. Thus, much of the difference between
observed and simulated trends is not caused by model errors, but rather by
random fluctuations of the climate system. This fact has strong implications
for the evaluation of simulated trends
<xref ref-type="bibr" rid="bib1.bibx1 bib1.bibx30 bib1.bibx12" id="paren.21"/>, but it is also the reason why
cross-validation of bias correction fails in this context.</p>
      <p id="d1e254">As any cross-validation consists of repeated holdout evaluations, we will in
the following only consider the holdout method. In
Sect. <xref ref-type="sec" rid="Ch1.S5"/> we will discuss how the following results
generalize to a full cross-validation.</p>
</sec>
<sec id="Ch1.S3">
  <title>Analytical derivation</title>
      <p id="d1e265">Consider a simulated time series <inline-formula><mml:math id="M5" display="inline"><mml:mrow><mml:msub><mml:mi>x</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula> and an observed time series <inline-formula><mml:math id="M6" display="inline"><mml:mrow><mml:msub><mml:mi>y</mml:mi><mml:mi>i</mml:mi></mml:msub></mml:mrow></mml:math></inline-formula>.
Assume that an evaluation addresses the representation of some statistic such
as the long-term mean. Over the calibration period, we denote the simulated
and observed means as <inline-formula><mml:math id="M7" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M8" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula>,
respectively. Correspondingly, we denote them as <inline-formula><mml:math id="M9" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> and
<inline-formula><mml:math id="M10" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> over the validation period. Then an estimate for the
bias over the calibration period is given as
          <disp-formula id="Ch1.E1" content-type="numbered"><mml:math id="M11" display="block"><mml:mrow><mml:mtext>BIAS</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula>
        Applying the bias estimate to the validation period, one obtains an estimate
of the corrected mean over the validation period:
          <disp-formula id="Ch1.E2" content-type="numbered"><mml:math id="M12" display="block"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext><mml:mtext>corr</mml:mtext></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>-</mml:mo><mml:mtext>BIAS</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <?pagebreak page4869?><p id="d1e439">The remaining residual bias is then
          <disp-formula id="Ch1.E3" content-type="numbered"><mml:math id="M13" display="block"><mml:mrow><mml:msub><mml:mtext>BIAS</mml:mtext><mml:mtext>res</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext><mml:mtext>corr</mml:mtext></mml:msubsup><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>+</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e517">This residual bias can be expressed in terms of the observed and
simulated climate change signals. The change signal from calibration
to validation period is defined as
          <disp-formula id="Ch1.E4" content-type="numbered"><mml:math id="M14" display="block"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub></mml:mrow></mml:math></disp-formula>
        for the model and
          <disp-formula id="Ch1.E5" content-type="numbered"><mml:math id="M15" display="block"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>-</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub></mml:mrow></mml:math></disp-formula>
        for the observations. Thus, the residual bias is given as
          <disp-formula id="Ch1.E6" content-type="numbered"><mml:math id="M16" display="block"><mml:mrow><mml:msub><mml:mtext>BIAS</mml:mtext><mml:mtext>res</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>x</mml:mi><mml:mo>-</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>y</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e604">For variables such as precipitation, one often considers relative
changes. Here a corresponding derivation holds. The relative error is
defined as
          <disp-formula id="Ch1.E7" content-type="numbered"><mml:math id="M17" display="block"><mml:mrow><mml:mtext>RE</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>,</mml:mo></mml:mrow></mml:math></disp-formula>
        and the corrected mean over the validation period is given as
          <disp-formula id="Ch1.E8" content-type="numbered"><mml:math id="M18" display="block"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext><mml:mtext>corr</mml:mtext></mml:msubsup><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>/</mml:mo><mml:mtext>RE</mml:mtext><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>⋅</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e700">The residual relative error results in
          <disp-formula id="Ch1.E9" content-type="numbered"><mml:math id="M19" display="block"><mml:mrow><mml:msub><mml:mtext>RE</mml:mtext><mml:mtext>res</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:msubsup><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext><mml:mtext>corr</mml:mtext></mml:msubsup><mml:mo>/</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:mfrac style="display"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>⋅</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub><mml:mo>⋅</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub></mml:mrow></mml:mfrac></mml:mstyle><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e779">The relative change signal from calibration to validation period is defined as
          <disp-formula id="Ch1.E10" content-type="numbered"><mml:math id="M20" display="block"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub></mml:mrow></mml:math></disp-formula>
        for the model and
          <disp-formula id="Ch1.E11" content-type="numbered"><mml:math id="M21" display="block"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>y</mml:mi><mml:mo>=</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub><mml:mo>/</mml:mo><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>cal</mml:mtext></mml:msub></mml:mrow></mml:math></disp-formula>
        for the observations. Hence, the residual relative error is
          <disp-formula id="Ch1.E12" content-type="numbered"><mml:math id="M22" display="block"><mml:mrow><mml:msub><mml:mtext>RE</mml:mtext><mml:mtext>res</mml:mtext></mml:msub><mml:mo>=</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>x</mml:mi><mml:mo>/</mml:mo><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>y</mml:mi><mml:mo>.</mml:mo></mml:mrow></mml:math></disp-formula></p>
      <p id="d1e866">The residual bias or relative error could further be tested for significance,
i.e. whether the bias-corrected statistic <inline-formula><mml:math id="M23" display="inline"><mml:mrow><mml:msubsup><mml:mover accent="true"><mml:mi>x</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext><mml:mtext>corr</mml:mtext></mml:msubsup></mml:mrow></mml:math></inline-formula> is significantly different from the observed
statistic <inline-formula><mml:math id="M24" display="inline"><mml:mrow><mml:msub><mml:mover accent="true"><mml:mi>y</mml:mi><mml:mo mathvariant="normal">‾</mml:mo></mml:mover><mml:mtext>val</mml:mtext></mml:msub></mml:mrow></mml:math></inline-formula> over the validation period. Thus, a
holdout evaluation will yield a positive result (in the sense of rejecting
the null hypothesis, i.e. a non-zero residual bias) if the simulated change
<inline-formula><mml:math id="M25" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:math></inline-formula> is different from the observed change <inline-formula><mml:math id="M26" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:math></inline-formula>, and a negative
result (i.e. a residual bias compatible with zero) if simulated and observed
changes are indistinguishable.</p>
      <p id="d1e919">Assume now that a given bias correction may or may not be sensible. Note in
this context that it is completely irrelevant to explicitly define what
constitutes a sensible bias correction (but for a brief discussion see
Sect. <xref ref-type="sec" rid="Ch1.S4"/>). Thus, in principle four cases are possible.
<list list-type="order"><list-item>
      <p id="d1e926">True negative: the bias correction is sensible, and the (bias-corrected)
climate model simulates a trend closely resembling the observed
trend.</p></list-item><list-item>
      <p id="d1e930">False positive: the bias correction is sensible, but due to internal climate
variability, the (bias-corrected) climate model simulates a trend
different from the observed trend.</p></list-item><list-item>
      <p id="d1e934">False negative: the bias correction is not sensible, but the (bias-corrected)
climate model for some reason simulates a trend similar to the
observed trend. This case corresponds to the example given in
<xref ref-type="bibr" rid="bib1.bibx21" id="normal.22"/>.</p></list-item><list-item>
      <p id="d1e941">True positive: the bias correction is not sensible, and the
(bias-corrected) climate model simulates a trend different from the
observed trend.</p></list-item></list></p>
      <p id="d1e944">The crucial point is that for typical record lengths, much of the difference
between simulated and observed changes <inline-formula><mml:math id="M27" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>x</mml:mi></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M28" display="inline"><mml:mrow><mml:mi mathvariant="normal">Δ</mml:mi><mml:mi>y</mml:mi></mml:mrow></mml:math></inline-formula> will be
caused by internal climate variability. Thus the result of a
cross-validation, i.e. which of the four cases occurs, is mostly random and
says only little about the sensibility of the cross-validation.</p>
      <p id="d1e967"><xref ref-type="bibr" rid="bib1.bibx21" id="normal.23"/> considered case 3: as the difference between simulated
and observed trends on typical timescales of a few decades is dominated by
internal variability, the holdout method is not suitable to identify a
non-sensible bias correction. The reverse conclusion is that the holdout
method – and consequently also a cross-validation – is not able to
corroborate whether a bias correction is sensible.</p>
      <p id="d1e973">Yet the discussion above implies an even stronger conclusion: because case
2 might
randomly occur, a sensible bias correction may be rejected by a
cross-validation. Thus, even more importantly, cross-validation in the given
context is not just useless, but even misleading.</p>
</sec>
<sec id="Ch1.S4">
  <title>Empirical demonstration</title>

      <?xmltex \floatpos{t}?><fig id="Ch1.F1" specific-use="star"><caption><p id="d1e984"> Maps of relative changes in boreal mean
summer (JJA) precipitation, 1981–2005 relative to 1956–1980.
<bold>(a)</bold> EC-EARTH, <bold>(b)</bold> E-OBS, version 15.0. Square: case 1 (true
negative); circle: case 2 (false positive); triangle: case 3
(false negative); diamond: case 4 (true negative).</p></caption>
        <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://hess.copernicus.org/articles/22/4867/2018/hess-22-4867-2018-f01.pdf"/>

      </fig>

      <p id="d1e999">To further illustrate the analytic findings, we will give examples of the
four cases in an exaggerated modelling example. We consider mean summer (JJA)
precipitation at four locations. As observational reference we select the
E-OBS data set <xref ref-type="bibr" rid="bib1.bibx9" id="paren.24"/>. As calibration period we use 1956–1980, as
evaluation period 1981–2005.</p>
      <p id="d1e1005">We need to select two examples where the given bias correction is sensible,
and two where it is not. Finding a convincing example of a sensible bias
correction has to rely on process understanding <xref ref-type="bibr" rid="bib1.bibx21" id="paren.25"/>. A major
precondition is that the climate model simulates a realistic present climate
and a credible climate change <xref ref-type="bibr" rid="bib1.bibx19" id="paren.26"/>. The former condition mainly
involves a realistic representation of the large-scale circulation
<xref ref-type="bibr" rid="bib1.bibx21" id="paren.27"/>. We therefore consider the following set-up: as<?pagebreak page4870?> examples of
a sensible bias correction, we consider summer mean precipitation at two
locations in Norway. Summer mean precipitation in Norway is dominated by
large-scale precipitation, which can sensibly be assumed to be realistically
simulated by current-generation general circulation models (GCMs).
Specifically we choose a transient simulation of EC-EARTH
<xref ref-type="bibr" rid="bib1.bibx10" id="paren.28"/>, a model which has been demonstrated to suffer from minor
biases in the synoptic-scale atmospheric circulation over Europe only
<xref ref-type="bibr" rid="bib1.bibx33" id="paren.29"/>. We assume that other potential problems such as mislocations
<xref ref-type="bibr" rid="bib1.bibx18" id="paren.30"/> or scale gaps <xref ref-type="bibr" rid="bib1.bibx15" id="paren.31"/> are negligible for the
considered locations and timescales. In this setting, we argue that a bias
correction is in principle sensible. Two slightly different locations have
been selected (Fig. <xref ref-type="fig" rid="Ch1.F1"/>): in the Børgefjell region north-east of
Trondheim observed and simulated trends are very similar (case 1). Further
north, around the town of Bodø, the two trends are very different as
observed precipitation has decreased whereas simulated precipitation has
increased. It is reasonable to assume that the observed negative trend is
caused by internal climate variability and not a forced change. Thus, in
principle the bias correction is sensible even though the validation rejects
it (case 2). Here, the key point is not whether the bias correction in this
particular example is really sensible or not, but that case 2 is indeed a
possible – and misleading – outcome of a bias correction.</p>
      <p id="d1e1032">In the following we show two examples where a bias correction is not
sensible. A discussion about the question when a bias correction makes no
sense would go very much beyond the scope of this piece. Therefore, we follow
the logic of <xref ref-type="bibr" rid="bib1.bibx21" id="normal.32"/> and select examples where model simulation and
observation are taken from geographically far away and climatically rather
different regions. The underlying idea is that for such cases, the model does
not represent the target variable such that a bias correction is without
doubt not sensible. Specifically, we consider the following two cases (see
Fig. <xref ref-type="fig" rid="Ch1.F1"/>): first, mapping simulated boreal summer mean
precipitation from the sub-tropical Maputo area (Mozambique, close to the
South African border) to the Taiga region of the Norwegian–Finnish boarder.
Here, observed and simulated trends are randomly similar. This example is
equivalent to that given in <xref ref-type="bibr" rid="bib1.bibx21" id="normal.33"/>, where a non-sensible bias
correction is not identified by the validation (case 3). Second, we map
summer mean precipitation from the tropical climate at Belén in the
Amazon delta to the mild and maritime climate of northern Portugal. Here,
observed and simulated trends are randomly very different: positive in the
Amazon delta, negative in Portugal. The non-sensible bias correction is thus correctly
identified (case 4). In these two examples, it was a priori obvious that a
bias correction is not sensible. In real applications, of course, such a
priori reasoning will be much more difficult and has to rest upon process
understanding (see discussion above). The key point, again, is that case 3 is
a possible outcome of a bias correction.</p>

      <?xmltex \floatpos{t}?><fig id="Ch1.F2" specific-use="star"><caption><p id="d1e1046"> Time series of boreal summer (JJA)
precipitation. <bold>(a)</bold> case 1 (true negative); <bold>(a)</bold> case 2 (false
positive); <bold>(c)</bold> case 3 (false negative); <bold>(d)</bold> case 4 (true
negative). Black: E-OBS; red: raw EC-EARTH; blue: bias-corrected
EC-EARTH. The straight horizontal lines depict the long-term means over
calibration and validation periods, respectively.</p></caption>
        <?xmltex \igopts{width=398.338583pt}?><graphic xlink:href="https://hess.copernicus.org/articles/22/4867/2018/hess-22-4867-2018-f02.pdf"/>

      </fig>

      <p id="d1e1067">Figure <xref ref-type="fig" rid="Ch1.F2"/> shows observed and simulated time series, the latter
before and after bias correction, for the four cases we considered. Figure 2a
and b show the sensible examples, Fig. 2c and d the non-sensible examples. In
Fig. 2a and c, observed and simulated trends randomly agree, and in Fig. 2b
and d they randomly disagree. As shown analytically in
Sect. <xref ref-type="sec" rid="Ch1.S3"/>, the residual bias vanishes in cases (a) and
(c), where the relative trends in observations and simulations
are similar, and it does not vanish in (b) and (d), where the relative trends
in observations and simulations disagree. The relevant cases are (b) and (c):
in the former, the bias correction is in principle sensible, but the holdout
method would suggest that it was not sensible (false positive). In the
latter, the bias correction is not sensible, but the holdout method would
suggest that it was sensible (false negative). These examples clearly
illustrate our previous reasoning: the holdout method, and thus<?pagebreak page4871?> also
cross-validation, yields misleading results when it is used to assess the
sensibility of bias-corrected climate change simulations against
observations.</p>
</sec>
<sec id="Ch1.S5" sec-type="conclusions">
  <title>Conclusions</title>
      <p id="d1e1081">We have demonstrated both analytically and with a modelling example that
cross-validation of free-running bias-corrected climate change simulations
against observations is misleading. The underlying reasoning is as follows:
the result of a cross-validation – a significant or non-significant residual
bias in the validation period – depends on the difference between observed
and simulated changes between calibration and validation periods. For typical
lengths of calibration and validation periods, these differences depend
mostly on the realizations of internal variability in the observations and
climate model. These differences therefore do not allow for conclusions about
the sensibility of a bias correction. As in any setting of significance
testing, four cases are possible: true negative, false positive, false
negative and true negative. The actual outcome in a given application is
mostly random.</p>
      <p id="d1e1084">The relevance of internal variability in the discussed cross-validation
context depends on the relative strength of internal variability and forced
trends, and the length of the calibration and validation period compared to
the periodicity of the dominant modes of internal climate variability. In
tropical climates, interannual variability such as that of El
Niño–Southern Oscillation dominates climate variability. Thus, relatively
short periods of a few decades may suffice to obtain stable estimates of
forced changes between calibration and validation period, and therefore to
assess whether a bias correction performs well given the observed changes. In
mid-latitude climates, however, the dominant modes of internal variability
have periodicities of several decades <xref ref-type="bibr" rid="bib1.bibx14 bib1.bibx26" id="paren.34"/> such
that stable estimates of forced changes will in general not be obtained for
the typical calibration and validation periods of a few decades only. The
strength of internal variability depends also on the chosen variable (e.g.
higher for precipitation than for temperature) and index (higher for extremes
than for long-term means). Spatial aggregation will reduce regional
short-term internal variability <xref ref-type="bibr" rid="bib1.bibx16" id="paren.35"/>, but will not affect
long-term modes of variability, which typically have coherent spatial
patterns.</p>
      <p id="d1e1093">We have derived these conclusions for the mean and the holdout method, where
the bias correction is calibrated against one part of the data and validated
against its complement. Yet the results can in principle be transferred to
other statistics such as variances or individual quantiles, and to a full
cross-validation. The residual mean bias, however, is always zero in a full
cross-validation, as long as the individual folds have the same length. The
reason is that changing the calibration and validation period changes the
sign of the residual bias. When averaging the residual bias across the
different folds, it cancels out. For the variance or similar statistics, the
outcome depends on the way the cross-validation is carried out: if the
residual bias is calculated for each fold separately and then averaged (as
suggested in the classical literature), the behaviour is as for the mean. If
the residual bias is calculated over a concatenated cross-validated time
series (as is typically done in the atmospheric sciences), the bias
correction in cases (b) and (d) will yield extremely high residual biases
(because the shift in the mean is not removed in the variance calculation).</p>
      <?pagebreak page4872?><p id="d1e1096">The consequence of these findings is that cross-validation should not
be used when evaluating bias correction of free-running climate
simulations against observations. In fact, a framework for evaluating
bias correction of climate simulations is still missing and not
trivial. As discussed in <xref ref-type="bibr" rid="bib1.bibx21" id="normal.36"/>, we propose to evaluate
non-calibrated temporal, spatial and process-based aspects of the
simulated time series. Whether a bias correction makes sense for
climate change projections depends in particular on the realism and
the credibility of the underlying climate model simulation. A
process-based evaluation is thus a key prerequisite for a successful
bias correction <xref ref-type="bibr" rid="bib1.bibx21" id="paren.37"/>.</p>
</sec>

      
      </body>
    <back><notes notes-type="dataavailability">

      <p id="d1e1109">E-OBS data are available from the ECA&amp;D website at
<uri>https://www.ecad.eu/</uri> (last access: 21 July 2017). EC-EARTH data are
available from the ESGF nodes accessible via
<uri>https://esgf.llnl.gov/nodes.html</uri> (last access: 17 September 2018).</p>
  </notes><notes notes-type="authorcontribution">

      <p id="d1e1121">DM and MW developed the idea for the study.
DM conducted the analytical derivations, carried out the analysis, and wrote
the manuscript. MW commented on the manuscript. DM and MW discussed the
results.</p>
  </notes><notes notes-type="competinginterests">

      <p id="d1e1127">The authors declare that they have no conflict of
interest.</p>
  </notes><ack><title>Acknowledgements</title><p id="d1e1133">This study has been inspired by discussions in EU COST Action ES1102
VALUE.<?xmltex \hack{\newline}?><?xmltex \hack{\newline}?> Edited by: Carlo De Michele
<?xmltex \hack{\newline}?> Reviewed by: Uwe Ehret and Seth McGinnis</p></ack><ref-list>
    <title>References</title>

      <ref id="bib1.bibx1"><label>Bhend and Whetton(2013)</label><mixed-citation>
Bhend, J. and Whetton, P.: Consistency of simulated and observed regional
changes in temperature, sea level pressure and precipitation, Clim. Change,
118, 799–810, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx2"><label>Deser et al.(2012)Deser, Knutti, Solomon, and Phillips</label><mixed-citation>
Deser, C., Knutti, R., Solomon, S., and Phillips, A.: Communication of the role
of natural variability in future North American climate, Nat. Clim.
Change, 2, 775–779, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx3"><label>Dosio and Paruolo(2011)</label><mixed-citation>Dosio, A. and Paruolo, P.: Bias correction of the ENSEMBLES high resolution
climate change projections for use by impact models: Evaluation on the
present climate, J. Geophys. Res. Atmos., 116, D16106, <ext-link xlink:href="https://doi.org/10.1029/2011JD015934" ext-link-type="DOI">10.1029/2011JD015934</ext-link>, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx4"><label>Efron and Gong(1983)</label><mixed-citation>
Efron, B. and Gong, G.: A leisurely look at the bootstrap, the jackknife, and
cross-validation, Am. Stat., 37, 36–48, 1983.</mixed-citation></ref>
      <ref id="bib1.bibx5"><label>Gangopadhyay et al.(2011)Gangopadhyay, Pruitt, Brekke, and
Raff</label><mixed-citation>
Gangopadhyay, S., Pruitt, T., Brekke, L., and Raff, D.: Hydrologic projections
for the Western United States, EOS, 92, 441–442, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx6"><label>Girvetz et al.(2013)Girvetz, Maurer, Duffy, Ruesch, Thrasher, and
Zganjar</label><mixed-citation>
Girvetz, E., Maurer, E., Duffy, P., Ruesch, A., Thrasher, B., and Zganjar, C.:
Making climate data relevant to decision making: the important details of
spatial and temporal downscaling, The World Bank, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx7"><label>Gudmundson et al.(2012)Gudmundson, Bremnes, Haugen, and
Engen-Skaugen</label><mixed-citation>Gudmundsson, L., Bremnes, J. B., Haugen, J. E., and Engen-Skaugen, T.: Technical Note:
Downscaling RCM precipitation to the station scale using statistical transformations – a
comparison of methods, Hydrol. Earth Syst. Sci., 16, 3383–3390, <ext-link xlink:href="https://doi.org/10.5194/hess-16-3383-2012" ext-link-type="DOI">10.5194/hess-16-3383-2012</ext-link>, 2012.</mixed-citation></ref>
      <ref id="bib1.bibx8"><label>Hagemann et al.(2013)Hagemann, Chen, Clark, Folwell, Gosling,
Haddeland, Hannasaki, Heinke, Ludwig, Voss, and Wiltshire</label><mixed-citation>Hagemann, S., Chen, C., Clark, D. B., Folwell, S., Gosling, S. N., Haddeland, I.,
Hanasaki, N., Heinke, J., Ludwig, F., Voss, F., and Wiltshire, A. J.: Climate change
impact on available water resources obtained using multiple global climate and
hydrology models, Earth Syst. Dynam., 4, 129–144, <ext-link xlink:href="https://doi.org/10.5194/esd-4-129-2013" ext-link-type="DOI">10.5194/esd-4-129-2013</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx9"><label>Haylock et al.(2008)Haylock, Hofstra, Klein Tank, Klok, Jones, and
New</label><mixed-citation>Haylock, M., Hofstra, N., Klein Tank, A., Klok, E., Jones, P., and New, M.:
A European daily high-resolution gridded data set of surface temperature and
precipitation for 1950–2006, J. Geophys. Res., 113, D20119, <ext-link xlink:href="https://doi.org/10.1029/2008JD010201" ext-link-type="DOI">10.1029/2008JD010201</ext-link>, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx10"><label>Hazeleger et al.(2010)</label><mixed-citation>
Hazeleger, W., Severijns, C., Semmler, T., Ştefănescu, S., Yang,
S., Wang, X., Wyser, K., Dutra, E., Baldasano, J., Bintanja, R., Bougeault,
P., Caballero, R., Ekman, A., Christensen, J., van den Hurk, B., Jimenez,
P., Jones, C., Kållberg, P., Koenigk, T., Mc Grath, R., Miranda, P.,
van Noije, T., Palmer, T., Parodi, J., Schmith, T., Selten, F., Storelvmo,
T., Sterl, A., Tapamo, H., Vancoppenolle, M., Viterbo, P., and Willen, U.:
EC-Earth: a seamless earth-system prediction approach in action, B.
Am. Meteorol. Soc., 91, 1357–1363, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx11"><label>Jolliffe and Stephenson(2003)</label><mixed-citation>
Jolliffe, I. and Stephenson, D., eds.: Forecast Verification: A Practitioner's
Guide in Atmospheric Science, Wiley, 2003.</mixed-citation></ref>
      <ref id="bib1.bibx12"><label>Laprise(2014)</label><mixed-citation>
Laprise, R.: Comment on “The added value to global model projections of
climate change by dynamical downscaling: A case study over the continental US
using the GISS-ModelE2 and WRF models” by Racherla et al., J. Geophys.
Res., 119, 3877–3881, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx13"><label>Li et al.(2010)Li, Sheffield, and Wood</label><mixed-citation>Li, H., Sheffield, J., and Wood, E.: Bias correction of monthly precipitation
and temperature fields from Intergovernmental Panel on Climate Change AR4
models using equidistant quantile matching., J. Geophys. Res., 115,
D10101, <ext-link xlink:href="https://doi.org/10.1029/2009JD012882" ext-link-type="DOI">10.1029/2009JD012882</ext-link>, 2010.</mixed-citation></ref>
      <ref id="bib1.bibx14"><label>Mantua et al.(1997)Mantua, Hare, Zhang, Wallace, and
Francis</label><mixed-citation>
Mantua, N. J., Hare, S., Zhang, Y., Wallace, J., and Francis, R.: A Pacific
interdecadal climate oscillation with impacts on salmon production, B.
Am. Meteorol. Soc., 78, 1069–1079, 1997.</mixed-citation></ref>
      <ref id="bib1.bibx15"><label>Maraun(2013a)</label><mixed-citation>
Maraun, D.: Bias Correction, Quantile Mapping and Downscaling: Revisiting the
Inflation Issue, J. Climate, 26, 2137–2143, 2013a.</mixed-citation></ref>
      <ref id="bib1.bibx16"><label>Maraun(2013b)</label><mixed-citation>Maraun, D.: When will trends in European mean and heavy daily precipitation
emerge?, Env. Res. Lett., 8, 014004, <ext-link xlink:href="https://doi.org/10.1088/1748-9326/8/1/014004" ext-link-type="DOI">10.1088/1748-9326/8/1/014004</ext-link>, 2013b.</mixed-citation></ref>
      <ref id="bib1.bibx17"><label>Maraun(2016)</label><mixed-citation>Maraun, D.: Bias Correcting Climate Change Simulations – a Critical Review,
Curr. Clim. Change Rep., 2, 211–220, <ext-link xlink:href="https://doi.org/10.1007/s40641-016-0050-x" ext-link-type="DOI">10.1007/s40641-016-0050-x</ext-link>, 2016.</mixed-citation></ref>
      <ref id="bib1.bibx18"><label>Maraun and Widmann(2015)</label><mixed-citation>Maraun, D. and Widmann, M.: The representation of location by a regional climate model in complex
terrain, Hydrol. Earth Syst. Sci., 19, 3449–3456, <ext-link xlink:href="https://doi.org/10.5194/hess-19-3449-2015" ext-link-type="DOI">10.5194/hess-19-3449-2015</ext-link>, 2015.</mixed-citation></ref>
      <ref id="bib1.bibx19"><label>Maraun and Widmann(2018)</label><mixed-citation>
Maraun, D. and Widmann, M.: Statistical Downscaling and Bias Correction for
Climate Research, Cambridge University Press, Cambridge, 2018.</mixed-citation></ref>
      <ref id="bib1.bibx20"><label>Maraun et al.(2015)Maraun, Widmann, Gutierrez, Kotlarski, Chandler,
Hertig, Wibig, Huth, and Wilcke</label><mixed-citation>
Maraun, D., Widmann, M., Gutierrez, J., Kotlarski, S., Chandler, R., Hertig,
E., Wibig, J., Huth, R., and Wilcke, R.: VALUE:<?pagebreak page4873?> A Framework to Validate
Downscaling Approaches for Climate Change Studies, Earth's Future, 3, 1–14,
2015.</mixed-citation></ref>
      <ref id="bib1.bibx21"><label>Maraun et al.(2017)Maraun, Shepherd, Widmann, Zappa, Walton, Hall,
Gutierrez, Hagemann, Richter, Soares, and Mearns</label><mixed-citation>Maraun, D., Shepherd, T., Widmann, M., Zappa, G., Walton, D., Hall, A.,
Gutierrez, J. M., Hagemann, S., Richter, I., Soares, P., and Mearns, L.:
Towards process-informed bias correction of climate change simulations, Nat.
Clim. Change, 7, 764–773, <ext-link xlink:href="https://doi.org/10.1038/NCLIMATE3418" ext-link-type="DOI">10.1038/NCLIMATE3418</ext-link>, 2017.</mixed-citation></ref>
      <ref id="bib1.bibx22"><label>Mason(2008)</label><mixed-citation>
Mason, S.: Understanding forecast verification statistics, Meteorol. Appl., 15,
31–40, 2008.</mixed-citation></ref>
      <ref id="bib1.bibx23"><label>Michaelsen(1987)</label><mixed-citation>
Michaelsen, J.: Cross-validation in statistical climate forecast models, J.
Clim. Appl. Meteorol., 26, 1589–1600, 1987.</mixed-citation></ref>
      <ref id="bib1.bibx24"><label>Piani et al.(2010a)Piani, Haerter, and
Coppola</label><mixed-citation>
Piani, C., Haerter, J., and Coppola, E.: Statistical bias correction for daily
precipitation in regional climate models over Europe, Theor. Appl.
Climatol., 99, 187–192, 2010a.</mixed-citation></ref>
      <ref id="bib1.bibx25"><label>Piani et al.(2010b)Piani, Weedon, Best, Gomes, Viterbo,
Hagemann, and Haerter</label><mixed-citation>
Piani, C., Weedon, G., Best, M., Gomes, S., Viterbo, P., Hagemann, S., and
Haerter, J.: Statistical bias correction of global simulated daily
precipitation and temperature for the application of hydrological models, J.
Hydrol., 395, 199–215, 2010b.</mixed-citation></ref>
      <ref id="bib1.bibx26"><label>Schlesinger and Ramankutty(1994)</label><mixed-citation>
Schlesinger, M. and Ramankutty, N.: An oscillation in the global climate system
of period 65–70 years, Nature, 367, 723–726, 1994.</mixed-citation></ref>
      <ref id="bib1.bibx27"><label>Stone(1974)</label><mixed-citation>
Stone, M.: Cross-validatory choice and assessment of statistical predictions,
J. Roy. Stat. Soc. B, 32, 111–147, 1974.</mixed-citation></ref>
      <ref id="bib1.bibx28"><label>Teutschbein and Seibert(2012)</label><mixed-citation>Teutschbein, C. and Seibert, J.: Bias correction of regional climate model
simulations for hydrological climate-change impact studies: Review and
evaluation of different methods, J. Hydrol., 456, 12–29, 2012.
 </mixed-citation></ref><?xmltex \hack{\newpage}?>
      <ref id="bib1.bibx29"><label>Themeßl et al.(2011)Themeßl, Gobiet, and
Leuprecht</label><mixed-citation>
Themeßl, M. J., Gobiet, A., and Leuprecht, A.: Empirical-statistical
downscaling and error correction of daily precipitation from regional climate
models, Int. J. Climatol., 31, 1530–1544, 2011.</mixed-citation></ref>
      <ref id="bib1.bibx30"><label>van Oldenborgh et al.(2013)van Oldenborgh, Doblas Reyes,
Drijfhout, and Hawkins</label><mixed-citation>van Oldenborgh, G., Doblas Reyes, F.-J., Drijfhout, S., and Hawkins, E.:
Reliability of regional climate model trends, Environ Res. Lett.,
8, 014055, <ext-link xlink:href="https://doi.org/10.1088/1748-9326/8/1/014055" ext-link-type="DOI">10.1088/1748-9326/8/1/014055</ext-link>, 2013.</mixed-citation></ref>
      <ref id="bib1.bibx31"><label>Warszawski et al.(2014)Warszawski, Frieler, Huber, Piontek,
Serdeczny, and Schewe</label><mixed-citation>Warszawski, L., Frieler, K., Huber, V., Piontek, F., Serdeczny, O., and Schewe,
J.: The Inter-Sectoral Impact Model Intercomparison Project (ISI–MIP):
Project framework, Proc. Nat. Acad. Sci., 111, 3228–3232,
<ext-link xlink:href="https://doi.org/10.1073/pnas.1312330110" ext-link-type="DOI">10.1073/pnas.1312330110</ext-link>, 2014.</mixed-citation></ref>
      <ref id="bib1.bibx32"><label>Wilks(2006)</label><mixed-citation>
Wilks, D. S.: Statistical Methods in the Atmospheric Sciences, Academic
Press/Elsevier, 2 Edn., 2006.</mixed-citation></ref>
      <ref id="bib1.bibx33"><label>Zappa et al.(2013)Zappa, Shaffrey, and Hodges</label><mixed-citation>
Zappa, G., Shaffrey, L., and Hodges, K.: The ability of CMIP5 models to
simulate North Atlantic extratropical cyclones, J. Climate, 26,
5379–5396, 2013.</mixed-citation></ref>

  </ref-list></back>
    <!--<article-title-html>Cross-validation of bias-corrected climate simulations is misleading</article-title-html>
<abstract-html><p>We demonstrate both analytically and with a modelling example that
cross-validation of free-running bias-corrected climate change simulations
against observations is misleading. The underlying reasoning is as follows: a
cross-validation can have in principle two outcomes. A negative (in the sense
of not rejecting a null hypothesis), if the residual bias in the validation
period after bias correction vanishes; and a positive, if the residual bias
in the validation period after bias correction is large. It can be shown
analytically that the residual bias depends solely on the difference between
the simulated and observed change between calibration and validation periods.
This change, however, depends mainly on the realizations of internal
variability in the observations and climate model. As a consequence, the
outcome of a cross-validation is also dominated by internal variability, and
does not allow for any conclusion about the sensibility of a bias correction.
In particular, a sensible bias correction may be rejected (false positive)
and a non-sensible bias correction may be accepted (false negative). We
therefore propose to avoid cross-validation when evaluating bias correction
of free-running bias-corrected climate change simulations against
observations. Instead, one should evaluate non-calibrated temporal, spatial
and process-based aspects.</p></abstract-html>
<ref-html id="bib1.bib1"><label>Bhend and Whetton(2013)</label><mixed-citation>
Bhend, J. and Whetton, P.: Consistency of simulated and observed regional
changes in temperature, sea level pressure and precipitation, Clim. Change,
118, 799–810, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib2"><label>Deser et al.(2012)Deser, Knutti, Solomon, and Phillips</label><mixed-citation>
Deser, C., Knutti, R., Solomon, S., and Phillips, A.: Communication of the role
of natural variability in future North American climate, Nat. Clim.
Change, 2, 775–779, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib3"><label>Dosio and Paruolo(2011)</label><mixed-citation>
Dosio, A. and Paruolo, P.: Bias correction of the ENSEMBLES high resolution
climate change projections for use by impact models: Evaluation on the
present climate, J. Geophys. Res. Atmos., 116, D16106, <a href="https://doi.org/10.1029/2011JD015934" target="_blank">https://doi.org/10.1029/2011JD015934</a>, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib4"><label>Efron and Gong(1983)</label><mixed-citation>
Efron, B. and Gong, G.: A leisurely look at the bootstrap, the jackknife, and
cross-validation, Am. Stat., 37, 36–48, 1983.
</mixed-citation></ref-html>
<ref-html id="bib1.bib5"><label>Gangopadhyay et al.(2011)Gangopadhyay, Pruitt, Brekke, and
Raff</label><mixed-citation>
Gangopadhyay, S., Pruitt, T., Brekke, L., and Raff, D.: Hydrologic projections
for the Western United States, EOS, 92, 441–442, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib6"><label>Girvetz et al.(2013)Girvetz, Maurer, Duffy, Ruesch, Thrasher, and
Zganjar</label><mixed-citation>
Girvetz, E., Maurer, E., Duffy, P., Ruesch, A., Thrasher, B., and Zganjar, C.:
Making climate data relevant to decision making: the important details of
spatial and temporal downscaling, The World Bank, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib7"><label>Gudmundson et al.(2012)Gudmundson, Bremnes, Haugen, and
Engen-Skaugen</label><mixed-citation>
Gudmundsson, L., Bremnes, J. B., Haugen, J. E., and Engen-Skaugen, T.: Technical Note:
Downscaling RCM precipitation to the station scale using statistical transformations – a
comparison of methods, Hydrol. Earth Syst. Sci., 16, 3383–3390, <a href="https://doi.org/10.5194/hess-16-3383-2012" target="_blank">https://doi.org/10.5194/hess-16-3383-2012</a>, 2012.
</mixed-citation></ref-html>
<ref-html id="bib1.bib8"><label>Hagemann et al.(2013)Hagemann, Chen, Clark, Folwell, Gosling,
Haddeland, Hannasaki, Heinke, Ludwig, Voss, and Wiltshire</label><mixed-citation>
Hagemann, S., Chen, C., Clark, D. B., Folwell, S., Gosling, S. N., Haddeland, I.,
Hanasaki, N., Heinke, J., Ludwig, F., Voss, F., and Wiltshire, A. J.: Climate change
impact on available water resources obtained using multiple global climate and
hydrology models, Earth Syst. Dynam., 4, 129–144, <a href="https://doi.org/10.5194/esd-4-129-2013" target="_blank">https://doi.org/10.5194/esd-4-129-2013</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib9"><label>Haylock et al.(2008)Haylock, Hofstra, Klein Tank, Klok, Jones, and
New</label><mixed-citation>
Haylock, M., Hofstra, N., Klein Tank, A., Klok, E., Jones, P., and New, M.:
A European daily high-resolution gridded data set of surface temperature and
precipitation for 1950–2006, J. Geophys. Res., 113, D20119, <a href="https://doi.org/10.1029/2008JD010201" target="_blank">https://doi.org/10.1029/2008JD010201</a>, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib10"><label>Hazeleger et al.(2010)</label><mixed-citation>
Hazeleger, W., Severijns, C., Semmler, T., Ştefănescu, S., Yang,
S., Wang, X., Wyser, K., Dutra, E., Baldasano, J., Bintanja, R., Bougeault,
P., Caballero, R., Ekman, A., Christensen, J., van den Hurk, B., Jimenez,
P., Jones, C., Kållberg, P., Koenigk, T., Mc Grath, R., Miranda, P.,
van Noije, T., Palmer, T., Parodi, J., Schmith, T., Selten, F., Storelvmo,
T., Sterl, A., Tapamo, H., Vancoppenolle, M., Viterbo, P., and Willen, U.:
EC-Earth: a seamless earth-system prediction approach in action, B.
Am. Meteorol. Soc., 91, 1357–1363, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib11"><label>Jolliffe and Stephenson(2003)</label><mixed-citation>
Jolliffe, I. and Stephenson, D., eds.: Forecast Verification: A Practitioner's
Guide in Atmospheric Science, Wiley, 2003.
</mixed-citation></ref-html>
<ref-html id="bib1.bib12"><label>Laprise(2014)</label><mixed-citation>
Laprise, R.: Comment on “The added value to global model projections of
climate change by dynamical downscaling: A case study over the continental US
using the GISS-ModelE2 and WRF models” by Racherla et al., J. Geophys.
Res., 119, 3877–3881, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib13"><label>Li et al.(2010)Li, Sheffield, and Wood</label><mixed-citation>
Li, H., Sheffield, J., and Wood, E.: Bias correction of monthly precipitation
and temperature fields from Intergovernmental Panel on Climate Change AR4
models using equidistant quantile matching., J. Geophys. Res., 115,
D10101, <a href="https://doi.org/10.1029/2009JD012882" target="_blank">https://doi.org/10.1029/2009JD012882</a>, 2010.
</mixed-citation></ref-html>
<ref-html id="bib1.bib14"><label>Mantua et al.(1997)Mantua, Hare, Zhang, Wallace, and
Francis</label><mixed-citation>
Mantua, N. J., Hare, S., Zhang, Y., Wallace, J., and Francis, R.: A Pacific
interdecadal climate oscillation with impacts on salmon production, B.
Am. Meteorol. Soc., 78, 1069–1079, 1997.
</mixed-citation></ref-html>
<ref-html id="bib1.bib15"><label>Maraun(2013a)</label><mixed-citation>
Maraun, D.: Bias Correction, Quantile Mapping and Downscaling: Revisiting the
Inflation Issue, J. Climate, 26, 2137–2143, 2013a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib16"><label>Maraun(2013b)</label><mixed-citation>
Maraun, D.: When will trends in European mean and heavy daily precipitation
emerge?, Env. Res. Lett., 8, 014004, <a href="https://doi.org/10.1088/1748-9326/8/1/014004" target="_blank">https://doi.org/10.1088/1748-9326/8/1/014004</a>, 2013b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib17"><label>Maraun(2016)</label><mixed-citation>
Maraun, D.: Bias Correcting Climate Change Simulations – a Critical Review,
Curr. Clim. Change Rep., 2, 211–220, <a href="https://doi.org/10.1007/s40641-016-0050-x" target="_blank">https://doi.org/10.1007/s40641-016-0050-x</a>, 2016.
</mixed-citation></ref-html>
<ref-html id="bib1.bib18"><label>Maraun and Widmann(2015)</label><mixed-citation>
Maraun, D. and Widmann, M.: The representation of location by a regional climate model in complex
terrain, Hydrol. Earth Syst. Sci., 19, 3449–3456, <a href="https://doi.org/10.5194/hess-19-3449-2015" target="_blank">https://doi.org/10.5194/hess-19-3449-2015</a>, 2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib19"><label>Maraun and Widmann(2018)</label><mixed-citation>
Maraun, D. and Widmann, M.: Statistical Downscaling and Bias Correction for
Climate Research, Cambridge University Press, Cambridge, 2018.
</mixed-citation></ref-html>
<ref-html id="bib1.bib20"><label>Maraun et al.(2015)Maraun, Widmann, Gutierrez, Kotlarski, Chandler,
Hertig, Wibig, Huth, and Wilcke</label><mixed-citation>
Maraun, D., Widmann, M., Gutierrez, J., Kotlarski, S., Chandler, R., Hertig,
E., Wibig, J., Huth, R., and Wilcke, R.: VALUE: A Framework to Validate
Downscaling Approaches for Climate Change Studies, Earth's Future, 3, 1–14,
2015.
</mixed-citation></ref-html>
<ref-html id="bib1.bib21"><label>Maraun et al.(2017)Maraun, Shepherd, Widmann, Zappa, Walton, Hall,
Gutierrez, Hagemann, Richter, Soares, and Mearns</label><mixed-citation>
Maraun, D., Shepherd, T., Widmann, M., Zappa, G., Walton, D., Hall, A.,
Gutierrez, J. M., Hagemann, S., Richter, I., Soares, P., and Mearns, L.:
Towards process-informed bias correction of climate change simulations, Nat.
Clim. Change, 7, 764–773, <a href="https://doi.org/10.1038/NCLIMATE3418" target="_blank">https://doi.org/10.1038/NCLIMATE3418</a>, 2017.
</mixed-citation></ref-html>
<ref-html id="bib1.bib22"><label>Mason(2008)</label><mixed-citation>
Mason, S.: Understanding forecast verification statistics, Meteorol. Appl., 15,
31–40, 2008.
</mixed-citation></ref-html>
<ref-html id="bib1.bib23"><label>Michaelsen(1987)</label><mixed-citation>
Michaelsen, J.: Cross-validation in statistical climate forecast models, J.
Clim. Appl. Meteorol., 26, 1589–1600, 1987.
</mixed-citation></ref-html>
<ref-html id="bib1.bib24"><label>Piani et al.(2010a)Piani, Haerter, and
Coppola</label><mixed-citation>
Piani, C., Haerter, J., and Coppola, E.: Statistical bias correction for daily
precipitation in regional climate models over Europe, Theor. Appl.
Climatol., 99, 187–192, 2010a.
</mixed-citation></ref-html>
<ref-html id="bib1.bib25"><label>Piani et al.(2010b)Piani, Weedon, Best, Gomes, Viterbo,
Hagemann, and Haerter</label><mixed-citation>
Piani, C., Weedon, G., Best, M., Gomes, S., Viterbo, P., Hagemann, S., and
Haerter, J.: Statistical bias correction of global simulated daily
precipitation and temperature for the application of hydrological models, J.
Hydrol., 395, 199–215, 2010b.
</mixed-citation></ref-html>
<ref-html id="bib1.bib26"><label>Schlesinger and Ramankutty(1994)</label><mixed-citation>
Schlesinger, M. and Ramankutty, N.: An oscillation in the global climate system
of period 65–70 years, Nature, 367, 723–726, 1994.
</mixed-citation></ref-html>
<ref-html id="bib1.bib27"><label>Stone(1974)</label><mixed-citation>
Stone, M.: Cross-validatory choice and assessment of statistical predictions,
J. Roy. Stat. Soc. B, 32, 111–147, 1974.
</mixed-citation></ref-html>
<ref-html id="bib1.bib28"><label>Teutschbein and Seibert(2012)</label><mixed-citation>
Teutschbein, C. and Seibert, J.: Bias correction of regional climate model
simulations for hydrological climate-change impact studies: Review and
evaluation of different methods, J. Hydrol., 456, 12–29, 2012.

</mixed-citation></ref-html>
<ref-html id="bib1.bib29"><label>Themeßl et al.(2011)Themeßl, Gobiet, and
Leuprecht</label><mixed-citation>
Themeßl, M. J., Gobiet, A., and Leuprecht, A.: Empirical-statistical
downscaling and error correction of daily precipitation from regional climate
models, Int. J. Climatol., 31, 1530–1544, 2011.
</mixed-citation></ref-html>
<ref-html id="bib1.bib30"><label>van Oldenborgh et al.(2013)van Oldenborgh, Doblas Reyes,
Drijfhout, and Hawkins</label><mixed-citation>
van Oldenborgh, G., Doblas Reyes, F.-J., Drijfhout, S., and Hawkins, E.:
Reliability of regional climate model trends, Environ Res. Lett.,
8, 014055, <a href="https://doi.org/10.1088/1748-9326/8/1/014055" target="_blank">https://doi.org/10.1088/1748-9326/8/1/014055</a>, 2013.
</mixed-citation></ref-html>
<ref-html id="bib1.bib31"><label>Warszawski et al.(2014)Warszawski, Frieler, Huber, Piontek,
Serdeczny, and Schewe</label><mixed-citation>
Warszawski, L., Frieler, K., Huber, V., Piontek, F., Serdeczny, O., and Schewe,
J.: The Inter-Sectoral Impact Model Intercomparison Project (ISI–MIP):
Project framework, Proc. Nat. Acad. Sci., 111, 3228–3232,
<a href="https://doi.org/10.1073/pnas.1312330110" target="_blank">https://doi.org/10.1073/pnas.1312330110</a>, 2014.
</mixed-citation></ref-html>
<ref-html id="bib1.bib32"><label>Wilks(2006)</label><mixed-citation>
Wilks, D. S.: Statistical Methods in the Atmospheric Sciences, Academic
Press/Elsevier, 2 Edn., 2006.
</mixed-citation></ref-html>
<ref-html id="bib1.bib33"><label>Zappa et al.(2013)Zappa, Shaffrey, and Hodges</label><mixed-citation>
Zappa, G., Shaffrey, L., and Hodges, K.: The ability of CMIP5 models to
simulate North Atlantic extratropical cyclones, J. Climate, 26,
5379–5396, 2013.
</mixed-citation></ref-html>--></article>
