Multi-version Tensor Completion for Time-delayed Spatio-temporal Data

Multi-version Tensor Completion for Time-delayed Spatio-temporal Data
复制标题

DOI:
10.24963/ijcai.2021/400
复制
发表时间:
2021-05
期刊:
--
影响因子:
--
通讯作者:
Cheng Qian;Nikos Kargas;Cao Xiao;Lucas Glass;N. Sidiropoulos;Jimeng Sun
Cheng Qian;Nikos Kargas;Cao Xiao;Lucas Glass;N. Sidiropoulos;Jimeng Sun
中科院分区:
其他
文献类型:
--
作者:
Cheng Qian;Nikos Kargas;Cao Xiao;Lucas Glass;N. Sidiropoulos;Jimeng Sun

文献摘要

相似文献

实际上,由于各种数据加载延迟,现实世界中时空数据通常不完整或不准确。例如,案例计数的位置 - 疾病时间张量可以在某些位置或疾病中具有多个延迟的时间切片的延迟更新。恢复输入张量的此类缺失或嘈杂(报告不足的)元素可以看作是广义张量的完成问题。现有的张量完成方法通常假定i)丢失元素是随机分布的,ii)每个张量元件的噪声为I.I.D。零均值。对于时空张量数据,这两个假设都可能违反。我们经常观察到多个版本的输入张量,其报告噪声水平不同。噪声的量可能是时间或位置依赖性的,因为将更多更新逐渐引入张量。我们将此类动态数据建模为具有额外张量模式以捕获数据更新的多次张量张量。我们提出了一个低级张量模型,以预测随着时间的推移更新。我们证明我们的方法可以准确预测许多现实世界张量的基础真相值。与最佳基线方法相比,我们获得的根均值率高达27.2%。最后,我们扩展了随着时间的推移跟踪张量数据的方法,从而导致大量的计算节省。
Real-world spatio-temporal data is often incomplete or inaccurate due to various data loading delays. For example, a location-disease-time tensor of case counts can have multiple delayed updates of recent temporal slices for some locations or diseases. Recovering such missing or noisy (under-reported) elements of the input tensor can be viewed as a generalized tensor completion problem. Existing tensor completion methods usually assume that i) missing elements are randomly distributed and ii) noise for each tensor element is i.i.d. zero-mean. Both assumptions can be violated for spatio-temporal tensor data. We often observe multiple versions of the input tensor with different under-reporting noise levels. The amount of noise can be time- or location-dependent as more updates are progressively introduced to the tensor. We model such dynamic data as a multi-version tensor with an extra tensor mode capturing the data updates. We propose a low-rank tensor model to predict the updates over time. We demonstrate that our method can accurately predict the ground-truth values of many real-world tensors. We obtain up to 27.2% lower root mean-squared-error compared to the best baseline method. Finally, we extend our method to track the tensor data over time, leading to significant computational savings.