Judicious setting of Dynamic Time Warping's window width allows more accurate classification of time series

Judicious setting of Dynamic Time Warping's window width allows more accurate classification of time series
复制标题

DOI:
10.1109/bigdata.2017.8258009
复制
发表时间:
2017-12
期刊:
2017 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Hoang Anh Dau;Diego Furtado Silva;F. Petitjean;G. Forestier;A. Bagnall;Eamonn J. Keogh
Hoang Anh Dau;Diego Furtado Silva;F. Petitjean;G. Forestier;A. Bagnall;Eamonn J. Keogh
中科院分区:
其他
文献类型:
--
作者:
Hoang Anh Dau;Diego Furtado Silva;F. Petitjean;G. Forestier;A. Bagnall;Eamonn J. Keogh

文献摘要

被引文献

相似文献

虽然基于动态时间规整(DTW)的最近邻分类算法被认为是时间序列分类的强大基线,但近年来有太多的算法声称能够在一般情况下提高其精度。这些提议中的许多想法牺牲了基于DTW的分类器提供的实现的简单性,而获得的收益相当有限。然而,很明显,有时即使是很小的改善也可以在重要的医疗或金融领域产生巨大的影响。在这项工作中,我们提出了一个意想不到的结论:在优化DTW的性能方面,一个没有得到重视的“低挂果”可以产生改进,使其成为更强大的基线,缩小更复杂方法的大部分或全部改进差距。我们表明,目前用于学习DTW的唯一参数-允许的最大翘曲量-的方法,可能会在小训练集上给出错误的答案。我们介绍了一种简单的方法来缓解小训练集问题,方法是创建合成样本来帮助学习参数。我们在UCR时间序列档案上评估了我们的想法,并在秋季分类中进行了案例研究,证明了我们的算法在分类精度上有了显著的提高。
While the Dynamic Time Warping (DTW) — based Nearest-Neighbor Classification algorithm is regarded as a strong baseline for time series classification, in recent years there has been a plethora of algorithms that have claimed to be able to improve upon its accuracy in the general case. Many of these proposed ideas sacrifice the simplicity of implementation that DTW-based classifiers offer for rather modest gains. Nevertheless, there are clearly times when even a small improvement could make a large difference in an important medical or financial domain. In this work, we make an unexpected claim; an underappreciated “low hanging fruit” in optimizing DTW's performance can produce improvements that make it an even stronger baseline, closing most or all the improvement gap of the more sophisticated methods. We show that the method currently used to learn DTW's only parameter, the maximum amount of warping allowed, is likely to give the wrong answer for small training sets. We introduce a simple method to mitigate the small training set issue by creating synthetic exemplars to help learn the parameter. We evaluate our ideas on the UCR Time Series Archive and a case study in fall classification, and demonstrate that our algorithm produces significant improvement in classification accuracy.