Comment on “Bayesian recursive parameter estimation for hydrologic models” by M. Thiemann, M. Trosset, H. Gupta, and S. Sorooshian

Comment on “Bayesian recursive parameter estimation for hydrologic models” by M. Thiemann, M. Trosset, H. Gupta, and S. Sorooshian
复制标题

DOI:
10.1029/2001wr001183
复制
发表时间:
2003-05
影响因子:
5.4
通讯作者:
K. Beven;P. Young
K. Beven;P. Young
中科院分区:
地球科学1区
文献类型:
--
作者:
K. Beven;P. Young

文献摘要

被引文献

相似文献

[1]我们认为,Thiemann等人[2001]的论文(以下简称TTGS)虽然在其分析的基本假设的限制范围内形式上是正确的,但在通常与水文模型参数相关的不确定性方面并没有得出合理的结果。此外,我们认为,在这些假设合理的地方,可能会有更有效的方法来实现同样的目的。实际上,与之相对的BARE方法和GLUE方法代表了解决问题的可能贝叶斯方法范围中的两个极端位置。例如,本文的结果表明,在TTGS实现的BARE算法中,所有的误差都被视为“测量误差”,并且得到的估计模型参数没有任何明显的不确定性。换句话说,从TTGS的结果来看,估计的模型似乎是系统的几乎确定的或真实的表示。另一方面,在GLUE中,在对多个(非无误差)模型的预测进行加权时隐式地处理误差源,而不需要对测量误差模型进行强有力的假设(模型实现之间确实可能有所不同)。然而,这两种极端情况都只是部分适当的:较好的方法是将输入误差、模型结构误差和输出中的实际测量误差的影响分开。TTGS论文提出的最关键的问题是模型结构是正确的隐含假设。特别是,TTGS报告的BARE参数估计有效地收敛到点估计,参数不确定性明显很小。这一结果肯定很难从任何实际角度来证明;事实上,在第一个模拟例子中(它涉及一个线性的,两项纳什级联(NC)模型,假设已知降雨量输入和白色高斯测量噪声),它肯定是不合理的。在这种情况下,加性噪声的存在意味着参数估计应该收敛于一个正态概率分布,其均值和协方差取决于已处理的样本数量。此外,所需的参数估计及其相关的协方差矩阵可以用比BARE计算效率高得多的方式进行估计,使用递归贝叶斯参数估计算法,其中只有分布的前两个矩(均值和协方差)顺序更新,因此根本不需要蒙特卡罗模拟。[3]例如,最优、精炼的工具变量(RIV)参数估计算法[Young, 1984]在无约束传递函数模型中产生递归更新的最大似然(ML)参数估计[Young, 2002a]。在这个例子中,NC是一个受约束的二阶传递函数(两个一阶元素是相同的),因此标准RIV估计在ML意义上是次优的。尽管如此,该算法清楚地识别出TF是二阶的(使用YIC准则[Young, 1990]),并且基于数据集末端获得的估计值对TF参数(真实值:0.8521和0.0219)的递归估计为0.8568(0.0124)和0.0214(0.0009)。这里括号中的数字是估计的标准误差和0.8521的特征值估计,基于两个估计的特征值的平均值(通常略有不同,因为对TF参数的估计是不受约束的)。如果需要完全的最优性,那么有必要使用更复杂的、受约束的ML数值优化过程(仍然比BARE简单得多),这产生了估计0.8511(0.0023)和0.0213(0.0007)。将BARE(真值25)转换为停留时间(或衰退系数)估计,数据集末端的无约束RIV估计为25.8766(+2.6574,2.2340),ML估计为24.8187(+0.4221,0.4093)。(这里给出的范围是因为特征值估计到停留时间估计的转换扭曲了误差分布。)显然,模型参数的不确定性在数据集的末尾仍然存在,这是任何有限的具有测量噪声的数据集所期望的。最后,与TTGS论文图3所示的BARE情况相反,模型预测输出的估计标准误差带包含了所有的建模误差,这是它应该做的[见Young, 2002a]。[4]值得注意的是,在实践中,NC示例中的约束估计的复杂性是不必要的,除非在特殊情况下,因为线性的,相同的元素,NC通常不被认为是美国地球物理联盟2003年版权明显非线性的良好TF模型。0043-1397/03/2001WR001183
[1] We believe that the paper by Thiemann et al. [2001] (hereinafter referred to as TTGS), while formally correct within the limitations of the underlying assumptions of their analysis, does not lead to sensible results in respect of the uncertainties normally associated with the parameters of hydrological models. Moreover, we feel that where these assumptions are reasonable, there may be more effective ways of achieving the same ends. In effect, the BARE methodology and the GLUE approach that they contrast it with, represent two extreme positions in a spectrum of possible Bayesian approaches to the problem. For instance, the results presented in the paper suggest that, in the BARE algorithm as implemented by TTGS, all of the error is treated as if it were ‘‘measurement error’’ and the estimated model parameters are obtained without any appreciable uncertainty. In other words, the estimated model appears, from the TTGS results, to be an almost deterministic, or true, representation of the system. In GLUE, on the other hand, the sources of error are treated implicitly in weighting the predictions of multiple (non-error free) models, without strong assumptions about a measurement error model (which might indeed vary from model realization to model realization). However, both of these extremes are only partially adequate: a preferable approach would be to separate out the effects of errors in the inputs, the errors in the model structures and real measurement errors in the outputs. [2] The most critical issue raised by the TTGS paper is the implicit assumption that the model structure is correct. In particular, the BARE parameter estimates, as reported by TTGS, effectively converge to point estimates with apparently little parameter uncertainty. This result is surely difficult to justify in any practical terms; indeed, it certainly cannot be justified in the first simulation example (which involves a linear, two term Nash Cascade (NC) model with assumed known rainfall input and white, Gaussian measurement noise). In this case, the presence of additive noise means that the parameter estimates should converge to a normal probability distribution with mean and covariance that depend on the number of samples that have been processed. Moreover, the required parameter estimates and their associated covariance matrix can be estimated in a much more computationally efficient manner than in BARE, using a recursive Bayesian parameter estimation algorithm in which only the first two moments of the distribution (mean and covariance) are updated sequentially, so that no Monte Carlo simulation is necessary at all. [3] For example, the optimal, refined instrumental variable (RIV) parameter estimation algorithm [Young, 1984] yields recursively updated Maximum Likelihood (ML) estimates of the parameters in an unconstrained transfer function model [Young, 2002a]. In this example, the NC is a second order transfer function that is constrained (the two first order elements are identical), so the standard RIV estimates are sub-optimal in an ML sense. Even so, the algorithm clearly identifies that the TF is second order (using the YIC criterion [Young, 1990]) and the recursive estimates of the TF parameters (true values: 0.8521 and 0.0219) based on the estimates obtained at the end of the data set are 0.8568(0.0124) and 0.0214(0.0009). Here the figures in brackets are the estimated standard errors and the estimate of the eigenvalue at 0.8521, is based on the mean of the two estimated eigenvalues (normally slightly different because the estimates of the TF parameters are unconstrained). If complete optimality is required, then it is necessary to use a more complex, constrained ML numerical optimization procedure (still much simpler than BARE) and this yields estimates 0.8511(0.0023) and 0.0213(0.0007). Converted to the residence time (or recession coefficient) estimate, as obtained by BARE (true value 25), the unconstrained RIV estimates at the end of the data set are 25.8766(+2.6574, 2.2340) and the ML estimates are 24.8187(+0.4221, 0.4093). (The range is given here because the transformation of the eigenvalue estimate to the residence time estimated distorts the error distribution.) Clearly, uncertainty in the model parameters still exists at the end of the data set and this is what would be expected for any finite set of data with measurement noise. Finally, in contrast to the BARE situation shown in figure 3 of the TTGS paper, the estimated standard error band on the model predicted output encompasses all of the modeling errors, as it should do [see Young, 2002a]. [4] It is worth noting that, in practical terms, the complication of constrained estimation in the NC example would not be necessary, except in exceptional circumstances, because the linear, identical element, NC is not normally considered a good TF model of the obviously nonlinear Copyright 2003 by the American Geophysical Union. 0043-1397/03/2001WR001183