Comparing Forecast Skill

Comparing Forecast Skill
复制标题

DOI:
10.1175/mwr-d-14-00045.1
复制
发表时间:
2014-12-01
影响因子:
3.2
通讯作者:
Tippett, Michael K.
Tippett, Michael K.
中科院分区:
地球科学2区
文献类型:
--
作者:
DelSole, Timothy;Tippett, Michael K.

文献摘要

被引文献

相似文献

预测的一个基本问题是一个预测系统是否比另一个预测系统更熟练。如果技能是在公共时期或使用一组公共观察值计算的,一些常用的统计显着性检验无法正确回答这个问题,因为这些检验没有考虑样本技能估计之间的相关性。此外,这些测试的结果偏向于表明技能没有差异,这一事实对预测的改进具有重要影响。本文表明,偏差的大小由一些参数来表征,例如样本大小以及预测与其误差之间的相关性,令人惊讶的是,这些参数可以根据数据进行估计。对于典型的季节性预测来说,偏差很大,这意味着熟悉的测试可能会错误地判断季节性预测技能的差异微不足道。审查了适合评估同一时期内技能差异的四项测试。这些检验基于符号检验、Wilcoxon 符号秩检验、Morgan-Granger-Newbold 检验和排列检验。这些技术应用于北美多模式集合的 ENSO 后报,并表明气候预报系统第 2 版和加拿大气候模式第 3 版 (CanCM3) 优于其他模型,因为它们的平方误差更频繁地小于其他单一模型的平方误差。应该认识到,虽然某些模型在特定时期和变量的某种意义上可能更优越,但预测的组合通常比单独的单一模型更有技巧。事实上,多模型均值显着优于所有单一模型。
A basic question in forecasting is whether one prediction system is more skillful than another. Some commonly used statistical significance tests cannot answer this question correctly if the skills are computed on a common period or using a common set of observations, because these tests do not account for correlations between sample skill estimates. Furthermore, the results of these tests are biased toward indicating no difference in skill, a fact that has important consequences for forecast improvement. This paper shows that the magnitude of bias is characterized by a few parameters such as sample size and correlation between forecasts and their errors, which, surprisingly, can be estimated from data. The bias is substantial for typical seasonal forecasts, implying that familiar tests may wrongly judge that differences in seasonal forecast skill are insignificant. Four tests that are appropriate for assessing differences in skill over a common period are reviewed. These tests are based on the sign test, the Wilcoxon signed-rank test, the Morgan-Granger-Newbold test, and a permutation test. These techniques are applied to ENSO hindcasts from the North American Multimodel Ensemble and reveal that the Climate Forecast System, version 2, and the Canadian Climate Model, version 3 (CanCM3), outperform other models in the sense that their squared error is less than that of other single models more frequently. It should be recognized that while certain models may be superior in a certain sense for a particular period and variable, combinations of forecasts are often significantly more skillful than a single model alone. In fact, the multimodel mean significantly outperforms all single models.