Exact multi-length scale and mean invariant motif discovery

Exact multi-length scale and mean invariant motif discovery
复制标题

精确的多长度尺度和平均不变基序发现

DOI:
10.1007/s10489-015-0684-8
复制
发表时间:
2016
影响因子:
5.3
通讯作者:
Yasser Mohammad and Toyoaki Nishida
Yasser Mohammad and Toyoaki Nishida
中科院分区:
计算机科学2区
文献类型:
--
作者:
三浦惇貴・廣田雅春;野澤浩樹,横山昌平;徐美玲,赤羽克仁,佐藤誠;岡田孝;横井伯英;Yasser Mohammad and Toyoaki Nishida

文献摘要

相似文献

在时间序列中发现近似重复的模体(ARM)是数据挖掘中的一个活跃研究领域。精确模体发现被定义为有效地找到最相似的时间序列子序列对的问题,可以作为发现ARM的基础。解决这个问题最有效的算法是MK算法,它被设计用来在已知长度上找到具有最大相似度的单个时间序列对。本文对MK算法进行了三种扩展,使其能够同时使用欧几里德距离度量和尺度不变的归一化版本来寻找多个长度的topK相似的子序列。然后,将提出的算法应用于合成数据和真实世界数据,重点是在人类运动轨迹中发现手臂。
Discovering approximately recurrent motifs (ARMs) in timeseries is an active area of research in data mining. Exact motif discovery is defined as the problem of efficiently finding the most similar pairs of timeseries subsequences and can be used as a basis for discovering ARMs. The most efficient algorithm for solving this problem is the MK algorithm which was designed to find a single pair of timeseries subsequences with maximum similarity at a known length. This paper provides three extensions of the MK algorithm that allow it to find the topKsimilar subsequences at multiple lengths using both the Euclidean distance metric and scale invariant normalized version of it. The proposed algorithms are then applied to both synthetic data and real-world data with a focus on discovery of ARMs in human motion trajectories.