Making clustering in delay-vector space meaningful

Making clustering in delay-vector space meaningful
复制标题

DOI:
10.1007/s10115-006-0042-6
复制
发表时间:
2007-04-01
影响因子:
2.7
通讯作者:
Chen, Jason R.
Chen, Jason R.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Chen, Jason R.

文献摘要

被引文献

相似文献

序列时间序列聚类是一种用于从时间序列数据中提取重要特征的技术。该方法可以被证明是动态系统文献中使用的延迟向量空间形式主义中的聚类过程。最近,有一个令人震惊的说法是,顺序时间序列聚类是毫无意义的。这对文献中大量的工作产生了重要的影响,因为这样的说法使这些工作的贡献无效。在本文中,我们表明,顺序的时间序列聚类是没有意义的,在这些作品中突出的问题源于他们使用的欧氏距离度量的距离测量的延迟向量空间。作为一种解决方案,我们认为相当一般的一类时间序列,并提出了一个制度的基础上,可以存在于延迟向量之间的两种类型的相似性,自然产生一种替代的距离测量延迟向量空间中的欧几里得距离。我们表明,使用这种替代的距离度量,顺序时间序列聚类确实是有意义的。我们重复了一个关键的实验,在工作中的“无意义”的索赔的基础上,并表明我们的方法导致一个成功的聚类结果。
Sequential time series clustering is a technique used to extract important features from time series data. The method can be shown to be the process of clustering in the delay-vector space formalism used in the Dynamical Systems literature. Recently, the startling claim was made that sequential time series clustering is meaningless. This has important consequences for a significant amount of work in the literature, since such a claim invalidates these work's contribution. In this paper, we show that sequential time series clustering is not meaningless, and that the problem highlighted in these works stem from their use of the Euclidean distance metric as the distance measure in the delay-vector space. As a solution, we consider quite a general class of time series, and propose a regime based on two types of similarity that can exist between delay vectors, giving rise naturally to an alternative distance measure to Euclidean distance in the delay-vector space. We show that, using this alternative distance measure, sequential time series clustering can indeed be meaningful. We repeat a key experiment in the work on which the "meaningless" claim was based, and show that our method leads to a successful clustering outcome.