K Rotation-invariant similarity in time series using bag-of-patterns representation

K Rotation-invariant similarity in time series using bag-of-patterns representation
复制标题

DOI:
10.1007/s10844-012-0196-5
复制
发表时间:
2012-10-01
影响因子:
3.4
通讯作者:
Li, Yuan
Li, Yuan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Lin, Jessica;Khade, Rohan;Li, Yuan

文献摘要

被引文献

相似文献

十多年来,时间序列相似性搜索一直受到数据挖掘研究者的极大关注。因此,许多时间序列表示和距离措施已被提出。然而,大多数现有的工作时间序列相似性搜索依赖于基于形状的相似性匹配。虽然现有的一些方法对于短时间序列数据工作良好,但当序列较长时,它们通常无法产生令人满意的结果。对于长序列,更适合考虑基于高层结构的相似性。在这项工作中,我们提出了一个基于直方图的表示时间序列数据,类似于“袋的话”的方法,被广泛接受的文本挖掘和信息检索社区。我们进行了广泛的实验,并表明我们的方法优于领先的现有方法在聚类,分类和异常检测几十个真实的数据集。我们进一步证明了该表示允许在形状数据集中进行旋转不变的匹配。
For more than a decade, time series similarity search has been given a great deal of attention by data mining researchers. As a result, many time series representations and distance measures have been proposed. However, most existing work on time series similarity search relies on shape-based similarity matching. While some of the existing approaches work well for short time series data, they typically fail to produce satisfactory results when the sequence is long. For long sequences, it is more appropriate to consider the similarity based on the higher-level structures. In this work, we present a histogram-based representation for time series data, similar to the "bag of words" approach that is widely accepted by the text mining and information retrieval communities. We performed extensive experiments and show that our approach outperforms the leading existing methods in clustering, classification, and anomaly detection on dozens of real datasets. We further demonstrate that the representation allows rotation-invariant matching in shape datasets.