Time-Series Clustering Based on the Characterization of Segment Typologies

Time-Series Clustering Based on the Characterization of Segment Typologies
复制标题

DOI:
10.1109/tcyb.2019.2962584
复制
发表时间:
2021-11-01
影响因子:
11.8
通讯作者:
Hervas-Martinez, Cesar
Hervas-Martinez, Cesar
中科院分区:
计算机科学1区
文献类型:
--
作者:
Guijo-Rubio, David;Manuel Duran-Rosal, Antonio;Hervas-Martinez, Cesar

文献摘要

被引文献

相似文献

时间序列聚类是根据时间序列的相似性或特征对时间序列进行分组的过程。以前的方法通常结合时间序列的特定距离测量和标准聚类方法。然而,这些方法没有考虑每个时间序列的不同子序列的相似性,这可以用来更好地比较数据集的时间序列对象。在本文中,我们提出了一种由两个聚类阶段组成的时间序列聚类新技术。第一步,将最小二乘多项式分割过程应用于每个时间序列,该过程基于返回不同长度片段的增长窗口技术。然后,根据近似分段的模型系数和一组统计特征,将所有分段投影到相同的维度空间中。映射后,第一层次聚类阶段应用于所有映射的片段,返回每个时间序列的片段组。在定义另一个特定的映射过程之后,这些簇用于表示同一维空间中的所有时间序列。在第二个也是最后一个聚类阶段,所有时间序列对象都被分组。我们考虑内部聚类质量来自动调整算法的主要参数,这是分割的错误阈值。在 UCR 时间序列分类档案的 84 个数据集上获得的结果与三种最先进的方法进行了比较,表明该方法的性能非常有前途,尤其是在较大的数据集上。
Time-series clustering is the process of grouping time series with respect to their similarity or characteristics. Previous approaches usually combine a specific distance measure for time series and a standard clustering method. However, these approaches do not take the similarity of the different subsequences of each time series into account, which can be used to better compare the time-series objects of the dataset. In this article, we propose a novel technique of time-series clustering consisting of two clustering stages. In a first step, a least-squares polynomial segmentation procedure is applied to each time series, which is based on a growing window technique that returns different-length segments. Then, all of the segments are projected into the same dimensional space, based on the coefficients of the model that approximates the segment and a set of statistical features. After mapping, a first hierarchical clustering phase is applied to all mapped segments, returning groups of segments for each time series. These clusters are used to represent all time series in the same dimensional space, after defining another specific mapping process. In a second and final clustering stage, all the time-series objects are grouped. We consider internal clustering quality to automatically adjust the main parameter of the algorithm, which is an error threshold for the segmentation. The results obtained on 84 datasets from the UCR Time Series Classification Archive have been compared against three state-of-the-art methods, showing that the performance of this methodology is very promising, especially on larger datasets.