Temporal Dynamic Matrix Factorization for Missing Data Prediction in Large Scale Coevolving Time Series

Temporal Dynamic Matrix Factorization for Missing Data Prediction in Large Scale Coevolving Time Series
复制标题

用于大规模协同演化时间序列中缺失数据预测的时间动态矩阵分解

DOI:
10.1109/access.2016.2606242
复制
发表时间:
2016-09
期刊:
影响因子:
3.9
通讯作者:
Yufeng Chen
Yufeng Chen
中科院分区:
计算机科学3区
文献类型:
--
作者:
Weiwei Shi;Yongxin Zhu;Philip S. Yu;Tian Huang;Chang Wang;Yishu Mao;Yufeng Chen

文献摘要

参考文献

被引文献

相似文献

在实际应用中,时间序列数据的缺失是影响数据精确分析的一个重要因素。然而,大多数现有的方法可能是不可行的,或可能是低效的,以预测大规模的共同发展的时间序列中的缺失值。此外,时间序列的演化需要适当处理,以适应时间特性。此外,在许多领域产生的数据量比以往任何时候都要大。在本文中,我们已经采取了共同发展的时间序列的缺失数据预测的挑战,采用时间动态矩阵分解技术。首先,我们的方法被优化设计为在很大程度上利用每个时间序列的内部模式和跨多个源的时间序列的信息来构建初始模型。基于这一思想,我们引入了混合正则化项来约束矩阵分解的目标函数。然后,提出了时间动态矩阵分解,以有效地更新初始已训练的模型。在动态矩阵分解的过程中,采用批量更新和微调策略,建立了一个有效的和高效的模型。在真实数据集和合成数据集上的实验表明,该方法能有效地提高缺失数据预测的性能。即使当缺失率高达90%,我们提出的方法仍然表现出较低的预测误差。动态性能表明,该方法可以获得令人满意的效果和效率。此外,我们还演示了如何利用Apache Spark的高处理能力在大规模共同进化时间序列中执行缺失数据预测。
Data missing in collections of time series occurs frequently in practical applications and turns out to be a major menace to precise data analysis. However, most of the existing methods either might be infeasible or could be inefficient to predict the missing values in large-scale coevolving time series. Also, the evolving of time series needs to be handled properly to adapt to the temporal characteristic. Furthermore, more massive volume of data is generated in many areas than ever before. In this paper, we have taken up the challenge of missing data prediction in coevolving time series by employing temporal dynamic matrix factorization techniques. First, our approaches are optimally designed to largely utilize both the interior patterns of each time series and the information of time series across multiple sources to build an initial model. Based on the idea, we have imposed hybrid regularization terms to constrain the objective functions of matrix factorization. Then, temporal dynamic matrix factorization is proposed to effectively update the initial already trained models. In the process of dynamic matrix factorization, batch updating and fine-tuning strategies are also employed to build an effective and efficient model. Extensive experiments on real-world data sets and synthetic data set demonstrate that the proposed approaches can effectively improve the performance of missing data prediction. Even when the missing ratio reaches as high as 90%, our proposed methods still show low prediction errors. Dynamic performance demonstrates that the methods can obtain satisfactory effectiveness and efficiency. Furthermore, we have also demonstrated how to take advantage of the high processing power of Apache Spark to perform missing data prediction in large-scale coevolving time series.
DOI: 10.1007/s00034-015-0047-z
发表时间: 2015-04
期刊: Circuits, Systems, and Signal Processing
影响因子: --
作者:
Junyou Shi;Long Chen;Wei-Wei Cui-Wei
通讯作者: Junyou Shi;Long Chen;Wei-Wei Cui-Wei
DOI: 10.1109/jstars.2015.2424683
发表时间: 2015-10-01
影响因子: 5.5
作者:
Rathore, Muhammad Mazhar Ullah;Paul, Anand;Ji, Wen
通讯作者: Ji, Wen
DOI: 10.1007/bf02898786
发表时间: 1896-09
期刊: American Potato Journal
影响因子: --
作者:
K. Fernow
通讯作者: K. Fernow
通过考虑时间和空间依赖性,对交通流进行有效的缺失数据插补
DOI: 10.1016/j.trc.2013.05.008
发表时间: 2013-09-01
影响因子: 8.3
作者:
Li, Li;Li, Yuebiao;Li, Zhiheng
通讯作者: Li, Zhiheng
DOI: 10.1016/j.snb.2015.03.028
发表时间: 2015-08-01
影响因子: 8.4
作者:
Fonollosa, Jordi;Sheik, Sadique;Marco, Santiago
通讯作者: Marco, Santiago