Exploiting big data in time series forecasting: A cross-sectional approach

Exploiting big data in time series forecasting: A cross-sectional approach
复制标题

DOI:
10.1109/dsaa.2015.7344786
复制
发表时间:
2015-12
期刊:
2015 IEEE International Conference on Data Science and Advanced Analytics (DSAA)
影响因子:
--
通讯作者:
Claudio Hartmann;M. Hahmann;Wolfgang Lehner;Frank Rosenthal
Claudio Hartmann;M. Hahmann;Wolfgang Lehner;Frank Rosenthal
中科院分区:
其他
文献类型:
--
作者:
Claudio Hartmann;M. Hahmann;Wolfgang Lehner;Frank Rosenthal

文献摘要

被引文献

相似文献

预测时间序列数据是管理、规划和决策不可或缺的组成部分。随着大数据趋势的发展,在越来越多的应用领域中,大量的时间序列数据可以从众多的异质数据源中获取。这些领域的高度动态和经常波动的特点,再加上从各种来源收集这些数据的逻辑问题,给预测带来了新的挑战。传统的方法严重依赖广泛而完整的历史数据来建立时间序列模型,因此,如果时间序列很短,或者更重要的是,是间歇性的,则不再适用。此外,大量的时间序列必须在不同的聚合级别上进行预测,并且最好是低延迟,同时预测精度应该保持较高。这几乎是不可能的,因为保持传统的重点是为每个单独的时间序列创建一个预测模型。在本文中,我们提出了一种新的预测方法,称为横截面预测,以应对这些挑战。这种方法是专门为具有大量时间序列的大数据集设计的。我们的方法打破了现有的概念,只为整个时间序列创建了一个模型,并且只需要一小部分可用的数据来提供准确的预测。通过利用数据集所有时间序列中的可用数据,可以对缺失值进行补偿,并可以在任意聚集级别上快速计算出准确的预测结果。
Forecasting time series data is an integral component for management, planning and decision making. Following the Big Data trend, large amounts of time series data are available from many heterogeneous data sources in more and more applications domains. The highly dynamic and often fluctuating character of these domains in combination with the logistic problems of collecting such data from a variety of sources, imposes new challenges to forecasting. Traditional approaches heavily rely on extensive and complete historical data to build time series models and are thus no longer applicable if time series are short or, even more important, intermittent. In addition, large numbers of time series have to be forecasted on different aggregation levels with preferably low latency, while forecast accuracy should remain high. This is almost impossible, when keeping the traditional focus on creating one forecast model for each individual time series. In this paper we tackle these challenges by presenting a novel forecasting approach called cross-sectional forecasting. This method is especially designed for Big Data sets with a multitude of time series. Our approach breaks with existing concepts by creating only one model for a whole set of time series and requiring only a fraction of the available data to provide accurate forecasts. By utilizing available data from all time series of a data set, missing values can be compensated and accurate forecasting results can be calculated quickly on arbitrary aggregation levels.