Classification of time series by shapelet transformation

Classification of time series by shapelet transformation
复制标题

DOI:
10.1007/s10618-013-0322-1
复制
发表时间:
2014-07-01
影响因子:
4.8
通讯作者:
Bagnall, Anthony
Bagnall, Anthony
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hills, Jon;Lines, Jason;Bagnall, Anthony

文献摘要

被引文献

相似文献

时间序列分类(TSC)问题提出了一个具体的挑战,分类算法:如何衡量序列之间的相似性。shapelet是一个时间序列子序列,它允许基于局部、相位独立的形状相似性的TSC。基于Shapelet的分类使用Shapelet和系列之间的相似性作为区分特征。shapelet方法的一个好处是shapelet是可理解的,并且可以提供对问题域的洞察。原始的基于shapelet的分类器将shapelet发现算法嵌入到决策树中,并使用信息增益来评估候选者的质量,通过枚举搜索在树的每个节点上找到新的shapelet。随后的研究主要集中在加速搜索的技术上。我们研究如何最好地使用shapelet原语来构建分类器。我们提出了一个单扫描shapelet算法,找到最好的shapelet,这是用来产生一个转换后的数据集,其中每个功能代表的时间序列和shapelet之间的距离。相对于嵌入式方法的主要优点是,转换后的数据可以与任何分类器结合使用,并且不需要递归搜索形状。我们证明,转换后的数据,结合更复杂的分类器,提供更高的准确性比嵌入式shapelet树。我们还评估了三个相似性措施,产生等效的结果,在更短的时间内获得信息。最后,我们表明,通过进行变形后的聚类,我们可以提高转换后的数据的可解释性。我们在29个数据集上进行了实验:17个来自UCR存储库,12个我们自己提供。
Time-series classification (TSC) problems present a specific challenge for classification algorithms: how to measure similarity between series. A shapelet is a time-series subsequence that allows for TSC based on local, phase-independent similarity in shape. Shapelet-based classification uses the similarity between a shapelet and a series as a discriminatory feature. One benefit of the shapelet approach is that shapelets are comprehensible, and can offer insight into the problem domain. The original shapelet-based classifier embeds the shapelet-discovery algorithm in a decision tree, and uses information gain to assess the quality of candidates, finding a new shapelet at each node of the tree through an enumerative search. Subsequent research has focused mainly on techniques to speed up the search. We examine how best to use the shapelet primitive to construct classifiers. We propose a single-scan shapelet algorithm that finds the best shapelets, which are used to produce a transformed dataset, where each of the features represent the distance between a time series and a shapelet. The primary advantages over the embedded approach are that the transformed data can be used in conjunction with any classifier, and that there is no recursive search for shapelets. We demonstrate that the transformed data, in conjunction with more complex classifiers, gives greater accuracy than the embedded shapelet tree. We also evaluate three similarity measures that produce equivalent results to information gain in less time. Finally, we show that by conducting post-transform clustering of shapelets, we can enhance the interpretability of the transformed data. We conduct our experiments on 29 datasets: 17 from the UCR repository, and 12 we provide ourselves.