On the use of cross-validation for time series predictor evaluation

On the use of cross-validation for time series predictor evaluation
复制标题

DOI:
10.1016/j.ins.2011.12.028
复制
发表时间:
2012-05-15
影响因子:
8.1
通讯作者:
Benitez, Jose M.
Benitez, Jose M.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Bergmeir, Christoph;Benitez, Jose M.

文献摘要

被引文献

相似文献

在时间序列预测器评估中,我们观察到,在模型选择过程方面,一方面传统预测程序的评估与另一方面机器学习技术的评估之间存在差距。在传统的预测中,通常的做法是从每个时间序列的末尾保留一部分用于测试,并使用序列的其余部分进行训练。因此,它没有充分利用的数据,但理论问题的时间演变的影响和数据内的依赖关系,以及有关缺失值的实际问题被消除。另一方面,在评估用于时间序列预测的机器学习和其他回归方法时,通常使用交叉验证进行评估,很少注意到这些理论问题使交叉验证的基本假设无效。为了缩小这一差距,并研究在实践中不同的模型选择程序的后果,我们已经开发了一个严格的和广泛的实证研究。六种不同的模型选择程序,基于(i)交叉验证和(ii)使用系列的最后一部分进行评估,用于评估四种机器学习和其他回归技术在合成和真实世界时间序列上的性能。在我们的研究中没有发现理论缺陷的实际后果,但交叉验证技术的使用导致了更稳健的模型选择。为了利用“两全其美”,我们建议使用一个块的形式的交叉验证时间序列评价成为标准程序,从而利用所有可用的信息和规避的理论问题。(C)2012 Elsevier Inc. All rights reserved.
In time series predictor evaluation, we observe that with respect to the model selection procedure there is a gap between evaluation of traditional forecasting procedures, on the one hand, and evaluation of machine learning techniques on the other hand. In traditional forecasting, it is common practice to reserve a part from the end of each time series for testing, and to use the rest of the series for training. Thus it is not made full use of the data, but theoretical problems with respect to temporal evolutionary effects and dependencies within the data as well as practical problems regarding missing values are eliminated. On the other hand, when evaluating machine learning and other regression methods used for time series forecasting, often cross-validation is used for evaluation, paying little attention to the fact that those theoretical problems invalidate the fundamental assumptions of cross-validation. To close this gap and examine the consequences of different model selection procedures in practice, we have developed a rigorous and extensive empirical study. Six different model selection procedures, based on (i) cross-validation and (ii) evaluation using the series' last part, are used to assess the performance of four machine learning and other regression techniques on synthetic and real-world time series. No practical consequences of the theoretical flaws were found during our study, but the use of cross-validation techniques led to a more robust model selection. To make use of the "best of both worlds", we suggest that the use of a blocked form of cross-validation for time series evaluation became the standard procedure, thus using all available information and circumventing the theoretical problems. (C) 2012 Elsevier Inc. All rights reserved.