Error-bounded Approximate Time Series Joins using Compact Dictionary Representations of Time Series

Error-bounded Approximate Time Series Joins using Compact Dictionary Representations of Time Series
复制标题

使用时间序列的紧凑字典表示的有误差范围的近似时间序列连接

DOI:
--
复制
发表时间:
2021
期刊:
SDM
影响因子:
--
通讯作者:
Eamonn J. Keogh
Eamonn J. Keogh
中科院分区:
--
文献类型:
--
作者:
Chin;Yan;Junpeng Wang;Huiyuan Chen;Zhongfang Zhuang;Wei Zhang;Eamonn J. Keogh

文献摘要

参考文献

被引文献

相似文献

矩阵概要文件是一种有效的数据挖掘工具,它为时间序列数据提供了相似连接功能。矩阵概要文件的用户既可以使用内部相似度连接(即自连接)将时间序列与自身连接起来,也可以使用内部相似度连接将时间序列与另一个时间序列连接起来。通过调用其中一种或两种连接类型,矩阵概要文件可以帮助用户发现数据中的保守结构和异常结构。自五年前引入矩阵轮廓法以来,人们一直在努力提高近似连接的计算速度;然而,这些努力中的大多数只关注自连接。在这项工作中,我们证明了通过创建时间序列的紧凑“字典”表示,可以有效地执行带有错误有界保证的近似时间序列间相似性连接。使用字典表示代替原始时间序列,我们能够将异常挖掘系统的吞吐量提高至少20倍,而基本上没有降低准确性。作为副作用,字典还以语义上有意义的方式总结时间序列,并可以提供直观和可操作的见解。我们展示了基于字典的时间序列间相似性连接在医学和运输等不同领域的实用性。
The matrix profile is an effective data mining tool that provides similarity join functionality for time series data. Users of the matrix profile can either join a time series with itself using intra-similarity join (i.e., self-join) or join a time series with another time series using inter-similarity join. By invoking either or both types of joins, the matrix profile can help users discover both conserved and anomalous structures in the data. Since the introduction of the matrix profile five years ago, multiple efforts have been made to speed up the computation with approximate joins; however, the majority of these efforts only focus on self-joins. In this work, we show that it is possible to efficiently perform approximate inter-time series similarity joins with error bounded guarantees by creating a compact"dictionary"representation of time series. Using the dictionary representation instead of the original time series, we are able to improve the throughput of an anomaly mining system by at least 20X, with essentially no decrease in accuracy. As a side effect, the dictionaries also summarize the time series in a semantically meaningful way and can provide intuitive and actionable insights. We demonstrate the utility of our dictionary-based inter-time series similarity joins on domains as diverse as medicine and transportation.
MERLIN:海量时间序列档案中任意长度异常的无参数发现
DOI: 10.1109/icdm50108.2020.00147
发表时间: 2020
期刊: ICDM 2020
影响因子: --
作者:
Nakamura, Takaaki;Imamura, Makoto;Mercer, Ryan;Keogh, Eamonn
通讯作者: Keogh, Eamonn
矩阵配置文件 XVIII:使用学习的近似矩阵配置文件在快速移动的流中进行时间序列挖掘
DOI: 10.1109/icdm.2019.00104
发表时间: 2019
期刊: International Conference on Data Mining
影响因子: --
作者:
Zimmerman, Zachary;Shakibay Senobari, Nader;Funning, Gareth;Papalexakis, Evangelos;Oymak, Samet;Brisk, Philip;Keogh, Eamonn
通讯作者: Keogh, Eamonn