Optimizing dynamic time warping's window width for time series data mining applications

Optimizing dynamic time warping's window width for time series data mining applications
复制标题

DOI:
10.1007/s10618-018-0565-y
复制
发表时间:
2018-07-01
影响因子:
4.8
通讯作者:
Keogh, Eamonn
Keogh, Eamonn
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hoang Anh Dau;Silva, Diego Furtado;Keogh, Eamonn

文献摘要

被引文献

相似文献

动态时间规整(DTW)是一个非常有竞争力的距离测量大多数时间序列数据挖掘问题。从DTW获得最佳性能需要设置其唯一参数,即最大扭曲量(w)。在有大量数据的监督情况下,w通常在训练阶段通过交叉验证来设置。然而,这种方法对于小的训练集可能会产生次优的结果。对于无监督的情况,通过交叉验证进行学习是不可能的,因为我们无法访问标记数据。因此,许多从业者采取假设“越大越好”,并且他们使用计算资源所允许的w的最大值。然而,正如我们将展示的那样,在大多数情况下,这是一种产生劣质聚类的天真方法。此外,最佳扭曲窗口宽度通常在两个任务之间是不可转移的,即,对于单个数据集,从业者不能简单地将学习到的用于分类的最佳w应用于聚类,反之亦然。此外,我们还将证明,适当的扭曲量不仅取决于数据结构,还取决于数据集的大小。因此,即使从业者知道给定数据集的最佳设置,如果他们将该设置应用于该数据的更大版本,他们也可能会迷失方向。所有这些问题似乎在很大程度上不为人所知,或者至少在社区中不被重视。在这项工作中,我们证明了正确设置DTW的扭曲窗口宽度的重要性,我们还提出了新的方法来学习这个参数在监督和无监督的设置。我们提出的学习w的算法可以显著提高分类精度和聚类质量。我们证明了我们的新观察的正确性和我们的想法的效用,通过测试他们与100多个公开可用的数据集。我们强有力的结果使我们能够提出一个可能意想不到的主张;在优化DTW的性能方面,一个被低估的“低挂果实”可以产生改进,使其成为更强大的基线,缩小近年来提出的更复杂方法的大部分或所有改进差距。
Dynamic Time Warping (DTW) is a highly competitive distance measure for most time series data mining problems. Obtaining the best performance from DTW requires setting its only parameter, the maximum amount of warping (w). In the supervised case with ample data, w is typically set by cross-validation in the training stage. However, this method is likely to yield suboptimal results for small training sets. For the unsupervised case, learning via cross-validation is not possible because we do not have access to labeled data. Many practitioners have thus resorted to assuming that "the larger the better", and they use the largest value of w permitted by the computational resources. However, as we will show, in most circumstances, this is a na < ve approach that produces inferior clusterings. Moreover, the best warping window width is generally non-transferable between the two tasks, i.e., for a single dataset, practitioners cannot simply apply the best w learned for classification on clustering or vice versa. In addition, we will demonstrate that the appropriate amount of warping not only depends on the data structure, but also on the dataset size. Thus, even if a practitioner knows the best setting for a given dataset, they will likely be at a lost if they apply that setting on a bigger size version of that data. All these issues seem largely unknown or at least unappreciated in the community. In this work, we demonstrate the importance of setting DTW's warping window width correctly, and we also propose novel methods to learn this parameter in both supervised and unsupervised settings. The algorithms we propose to learn w can produce significant improvements in classification accuracy and clustering quality. We demonstrate the correctness of our novel observations and the utility of our ideas by testing them with more than one hundred publicly available datasets. Our forceful results allow us to make a perhaps unexpected claim; an underappreciated "low hanging fruit" in optimizing DTW's performance can produce improvements that make it an even stronger baseline, closing most or all the improvement gap of the more sophisticated methods proposed in recent years.