Parsimonious Linear Fingerprinting for Time Series

Parsimonious Linear Fingerprinting for Time Series
复制标题

DOI:
10.14778/1920841.1920893
复制
发表时间:
2010-09-01
影响因子:
2.5
通讯作者:
Faloutsos, Christos
Faloutsos, Christos
中科院分区:
计算机科学2区
文献类型:
--
作者:
Li, Lei;Prakash, B. Aditya;Faloutsos, Christos

文献摘要

被引文献

相似文献

研究了多时间序列的有效挖掘和汇总问题。我们提出了PLiF,一种新的方法来发现基本特征(“指纹”),通过利用联合动力学的数值序列。我们的指纹识别方法具有以下优点:(a)它导致可解释的特征;(B)它是通用的:PLiF能够实现许多挖掘任务,包括聚类、压缩、可视化、预测和分割,在每个任务中匹配顶级竞争对手;以及(c)它是快速且可扩展的,复杂度与序列长度呈线性关系。我们在合成和真实的数据集上进行了实验,包括人体动作捕捉数据(17 MB的人体动作)、传感器数据(166个传感器)和网络路由器流量数据(2年内1800万次原始更新)。尽管具有通用性,但PLiF在聚类方面优于顶级聚类方法;在压缩方面优于顶级压缩方法(对于相同的压缩比,重建误差是3倍);它提供了有意义的可视化,同时享有线性扩展。
We study the problem of mining and summarizing multiple time series effectively and efficiently. We propose PLiF, a novel method to discover essential characteristics ("fingerprints"), by exploiting the joint dynamics in numerical sequences. Our fingerprinting method has the following bene-fits: (a) it leads to interpretable features; (b) it is versatile: PLiF enables numerous mining tasks, including clustering, compression, visualization, forecasting, and segmentation, matching top competitors in each task; and (c) it is fast and scalable, with linear complexity on the length of the sequences.We did experiments on both synthetic and real datasets, including human motion capture data (17MB of human motions), sensor data (166 sensors), and network router traffic data (18 million raw updates over 2 years). Despite its generality, PLiF outperforms the top clustering methods on clustering; the top compression methods on compression (3 times better reconstruction error, for the same compression ratio); it gives meaningful visualization and at the same time, enjoys a linear scale-up.