Time series clustering in linear time complexity

Time series clustering in linear time complexity
复制标题

DOI:
10.1007/s10618-021-00798-w
复制
发表时间:
2021-09
影响因子:
4.8
通讯作者:
Xiaosheng Li;Jessica Lin;Liang Zhao
Xiaosheng Li;Jessica Lin;Liang Zhao
中科院分区:
计算机科学3区
文献类型:
--
作者:
Xiaosheng Li;Jessica Lin;Liang Zhao

文献摘要

相似文献

随着数据存储能力的增强以及数据生成和收集技术的进步,大量的时间序列数据变得可用,其内容也在迅速变化。这就要求数据挖掘方法具有较低的时间复杂度来处理庞大且快速变化的数据。本文提出了一种具有线性时间复杂度的时间序列聚类算法。该算法通过检查时间序列中随机选择的符号模式对数据进行分割。理论分析表明,这一过程可以揭示数据中的基团结构。我们对来自著名的UCR时间序列档案的所有128个数据集进行了广泛的评估,并与最先进的方法进行了统计分析比较。结果表明,与其他方法相比,该方法具有更好的精度。我们还进行了实验,探索算法的参数和配置如何影响最终的聚类结果。
With the increasing power of data storage and advances in data generation and collection technologies, large volumes of time series data become available and the content is changing rapidly. This requires data mining methods to have low time complexity to handle the huge and fast-changing data. This article presents a novel time series clustering algorithm that has linear time complexity. The proposed algorithm partitions the data by checking some randomly selected symbolic patterns in the time series. We provide theoretical analysis to show that group structures in the data can be revealed from this process. We evaluate the proposed algorithm extensively on all 128 datasets from the well-known UCR time series archive, and compare with the state-of-the-art approaches with statistical analysis. The results show that the proposed method achieves better accuracy compared with other rival methods. We also conduct experiments to explore how the parameters and configuration of the algorithm can affect the final clustering results.