Characteristic-based clustering for time series data

Characteristic-based clustering for time series data
复制标题

DOI:
10.1007/s10618-005-0039-x
复制
发表时间:
2006-11-01
影响因子:
4.8
通讯作者:
Hyndman, Rob
Hyndman, Rob
中科院分区:
计算机科学3区
文献类型:
--
作者:
Wang, Xiaozhe;Smith, Kate;Hyndman, Rob

文献摘要

被引文献

相似文献

随着时间序列聚类研究的日益重要,特别是对于医学或金融等长时间序列之间的相似性搜索,找到一种方法来解决大多数聚类方法在某些情况下不切实际的突出问题是至关重要的。当时间序列很长时,一些聚类算法可能会失败,因为相似度的标记在高维空间中是可疑的;当聚类基于距离度量时,许多方法无法处理丢失的数据。本文提出了一种基于时间序列结构特征的聚类方法。与其他替代方法不同,该方法不使用距离度量对点值进行聚类,而是基于从时间序列中提取的全局特征进行聚类。从每个单独的序列中获得特征度量,并可以馈送到任意的聚类算法中,包括无监督神经网络算法、自组织映射或分层聚类算法。描述时间序列的全局度量是通过应用统计操作获得的,这些统计操作最好地捕捉了潜在的特征:趋势、季节性、周期性、序列相关性、偏度、峰度、混沌、非线性和自相似性。由于该方法使用提取的全局度量进行聚类,因此降低了时间序列的维数,并且对丢失或有噪声的数据不太敏感。我们进一步提供了一种搜索机制,从特征集中找到应该用作聚类输入的最佳选择。所提出的技术已经使用之前报道的用于时间序列聚类的基准时间序列数据集和一组具有已知特征的时间序列数据集进行了测试。实证结果表明,我们的方法能够产生有意义的聚类。所得到的聚类与其他方法产生的聚类相似,但有一些有希望的和有趣的变化,可以用时间序列的全局特征的知识直观地解释。
With the growing importance of time series clustering research, particularly for similarity searches amongst long time series such as those arising in medicine or finance, it is critical for us to find a way to resolve the outstanding problems that make most clustering methods impractical under certain circumstances. When the time series is very long, some clustering algorithms may fail because the very notation of similarity is dubious in high dimension space; many methods cannot handle missing data when the clustering is based on a distance metric.This paper proposes a method for clustering of time series based on their structural characteristics. Unlike other alternatives, this method does not cluster point values using a distance metric, rather it clusters based on global features extracted from the time series. The feature measures are obtained from each individual series and can be fed into arbitrary clustering algorithms, including an unsupervised neural network algorithm, self-organizing map, or hierarchal clustering algorithm.Global measures describing the time series are obtained by applying statistical operations that best capture the underlying characteristics: trend, seasonality, periodicity, serial correlation, skewness, kurtosis, chaos, nonlinearity, and self-similarity. Since the method clusters using extracted global measures, it reduces the dimensionality of the time series and is much less sensitive to missing or noisy data. We further provide a search mechanism to find the best selection from the feature set that should be used as the clustering inputs.The proposed technique has been tested using benchmark time series datasets previously reported for time series clustering and a set of time series datasets with known characteristics. The empirical results show that our approach is able to yield meaningful clusters. The resulting clusters are similar to those produced by other methods, but with some promising and interesting variations that can be intuitively explained with knowledge of the global characteristics of the time series.