Clustering time series with clipped data

Clustering time series with clipped data
复制标题

DOI:
10.1007/s10994-005-5825-6
复制
发表时间:
2005-02-01
期刊:
影响因子:
7.5
通讯作者:
Janacek, G
Janacek, G
中科院分区:
计算机科学3区
文献类型:
--
作者:
Bagnall, A;Janacek, G

文献摘要

被引文献

相似文献

时间序列聚类是一个在各个领域都有应用的问题,并且最近吸引了大量的研究。时间序列数据通常很大并且可能包含异常值。我们证明了裁剪时间序列(离散到高于或低于中位数)的简单过程减少了内存需求,并显着加快了聚类速度,而不会降低聚类精度。我们还证明,当数据中存在异常值时,裁剪可以提高聚类准确性,从而作为异常值检测的手段和识别模型错误指定的方法。我们考虑多项式、自回归移动平均和隐马尔可夫模型的模拟数据,并表明聚类中使用的裁剪数据的估计参数渐近地趋向于未裁剪数据的估计参数。我们还通过实验证明,如果该系列足够长,则剪裁数据的准确性不会显着低于未剪裁数据的准确性,并且如果该系列包含异常值,则剪裁会导致明显更好的聚类。然后,我们说明使用剪裁序列如何在检测两个现实世界数据集(发电投标数据集和心电图数据集)上的模型错误指定和异常值方面发挥实际作用。
Clustering time series is a problem that has applications in a wide variety of fields, and has recently attracted a large amount of research. Time series data are often large and may contain outliers. We show that the simple procedure of clipping the time series (discretising to above or below the median) reduces memory requirements and significantly speeds up clustering without decreasing clustering accuracy. We also demonstrate that clipping increases clustering accuracy when there are outliers in the data, thus serving as a means of outlier detection and a method of identifying model misspecification. We consider simulated data from polynomial, autoregressive moving average and hidden Markov models and show that the estimated parameters of the clipped data used in clustering tend, asymptotically, to those of the unclipped data. We also demonstrate experimentally that, if the series are long enough, the accuracy on clipped data is not significantly less than the accuracy on unclipped data, and if the series contain outliers then clipping results in significantly better clusterings. We then illustrate how using clipped series can be of practical benefit in detecting model misspecification and outliers on two real world data sets: an electricity generation bid data set and an ECG data set.