Matrix Profile Index Approximation for Streaming Time Series

Matrix Profile Index Approximation for Streaming Time Series
复制标题

DOI:
10.1109/bigdata52589.2021.9671484
复制
发表时间:
2021-12
期刊:
2021 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Maryam Shahcheraghi;Trevor Cappon;Samet Oymak;E. Papalexakis;Eamonn J. Keogh;Zachary Zimmerman;P. Brisk
Maryam Shahcheraghi;Trevor Cappon;Samet Oymak;E. Papalexakis;Eamonn J. Keogh;Zachary Zimmerman;P. Brisk
中科院分区:
其他
文献类型:
--
作者:
Maryam Shahcheraghi;Trevor Cappon;Samet Oymak;E. Papalexakis;Eamonn J. Keogh;Zachary Zimmerman;P. Brisk

文献摘要

相似文献

在时间序列中发现图案(重复模式)是许多行业和科学领域的关键因素。时间序列,其中采样率不可能在计算上实时限制,并且强烈希望将计算尽可能接近传感器。机器学习模型为这些问题提供了近似的答案,例如,训练的模型可以预测最近采样的数据点窗口与用于培训的代表性时间序列。预测比赛的“强度”,而且还要在代表性时间序列中评估我们在两个不同的现实世界数据集上的方法。与精确的计算相比,40倍,预测精度高达87.9%,具体取决于预测的粒度。
Discovery of motifs (repeated patterns) in time series is a key factor across numerous industries and scientific fields. These and related problems have effectively been solved for offline analysis of time series; however, these approaches are computationally intensive and do not lend themselves to streaming time series, where the sampling rate imposes real-time constraints on computation and there is strong desire to locate computation as close as possible to the sensor. One promising solution is to use low-cost machine learning models to provide approximate answers to these problems. For example, prior work has trained models to predict the similarity of the most recently sampled window of data points to a representative time series used for training. This work addresses a more challenging problem: to predict not only the "strength" of the match, but also the relative location in the representative time series where the match occurs. We evaluate our approach on two different real world datasets; we demonstrate speedups as high as 40× compared to exact computations, with predictive accuracy as high as 87.9%, depending on the granularity of the prediction.