Significance of changes in medium-range forecast scores

Significance of changes in medium-range forecast scores
复制标题

中期预测分数变化的显着性

DOI:
10.3402/tellusa.v68.30229
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
A. Geer
A. Geer
中科院分区:
--
文献类型:
--
作者:
A. Geer

文献摘要

被引文献

相似文献

天气预报发展的影响是通过预报验证来衡量的,但许多发展虽然有用,但对中期预报分数的影响不到0.5%。 个体预测质量的混沌变化是如此之大,以至于在将这些“较小”的发展与对照进行比较时,很难达到统计学意义。例如,对于60个单独的预测,并且需要95%的置信水平,使用Student's t检验,第5天预测的质量变化需要大于1%才具有统计学显著性。  本研究的第一个目的是简单地说明在预测验证的重要性,并指出令人惊讶的大样本量,需要达到显着性。第二个目的是看看目前的显着性检验方法有多可靠,因为人们怀疑,表面上显着的结果实际上可能是由混乱的变异性产生的。零假设的独立实现可以使用包含纯数值扰动的预测实验来创建,并将其与控制进行比较。从大约2.5年的测试中得到1885个配对差异,可以构建一个替代的显著性检验,该检验对数据没有统计学假设。这是用来实验测试的预测分数的正常统计框架的有效性,它表明,学生的t检验的天真的应用程序确实产生了太多的假阳性(即假拒绝零假设)。一个已知的问题是预测分数的时间自相关性,这可以通过置信范围大小的通货膨胀来校正,但是典型的通货膨胀因子,例如基于AR(1)模型的通货膨胀因子,不够大,并且它们受到采样不确定性的影响。此外,统计多重性的重要性还没有得到重视,当许多实验放在一起比较时,这变得特别危险。例如,在三个预测实验中,可能有大约1/2的机会得到假阳性。然而,如果对自相关性进行了正确的调整,并且当多重性的影响使用Šidák校正进行了适当的处理时,t检验是发现预测分数变化的重要性的可靠方法。
The impact of developments in weather forecasting is measured using forecast verification, but many developments, though useful, have impacts of less than 0.5 % on medium-range forecast scores. Chaotic variability in the quality of individual forecasts is so large that it can be hard to achieve statistical significance when comparing these ‘smaller’ developments to a control. For example, with 60 separate forecasts and requiring a 95 % confidence level, a change in quality of the day-5 forecast needs to be larger than 1 % to be statistically significant using a Student's t-test. The first aim of this study is simply to illustrate the importance of significance testing in forecast verification, and to point out the surprisingly large sample sizes that are required to attain significance. The second aim is to see how reliable are current approaches to significance testing, following suspicion that apparently significant results may actually have been generated by chaotic variability. An independent realisation of the null hypothesis can be created using a forecast experiment containing a purely numerical perturbation, and comparing it to a control. With 1885 paired differences from about 2.5 yr of testing, an alternative significance test can be constructed that makes no statistical assumptions about the data. This is used to experimentally test the validity of the normal statistical framework for forecast scores, and it shows that the naive application of Student's t-test does generate too many false positives (i.e. false rejections of the null hypothesis). A known issue is temporal autocorrelation in forecast scores, which can be corrected by an inflation in the size of confidence range, but typical inflation factors, such as those based on an AR(1) model, are not big enough and they are affected by sampling uncertainty. Further, the importance of statistical multiplicity has not been appreciated, and this becomes particularly dangerous when many experiments are compared together. For example, across three forecast experiments, there could be roughly a 1 in 2 chance of getting a false positive. However, if correctly adjusted for autocorrelation, and when the effects of multiplicity are properly treated using a Šidák correction, the t-test is a reliable way of finding the significance of changes in forecast scores.