An autoencoder-based deep learning approach for clustering time series data

An autoencoder-based deep learning approach for clustering time series data
复制标题

DOI:
10.1007/s42452-020-2584-8
复制
发表时间:
2020-04
影响因子:
2.6
通讯作者:
Neda Tavakoli;Sima Siami‐Namini;Mahdi Adl Khanghah;Fahimeh Mirza Soltani;Akbar Siami Namin
Neda Tavakoli;Sima Siami‐Namini;Mahdi Adl Khanghah;Fahimeh Mirza Soltani;Akbar Siami Namin
中科院分区:
--
文献类型:
--
作者:
Neda Tavakoli;Sima Siami‐Namini;Mahdi Adl Khanghah;Fahimeh Mirza Soltani;Akbar Siami Namin

文献摘要

被引文献

相似文献

本文介绍了一种基于两阶段深度学习的时间序列数据聚类方法。首先,引入一种新的技术来利用这些特性(例如,波动性),以便创建标签,从而使问题能够从无监督学习转变为有监督学习。其次,构建基于自动编码器的深度学习模型,以建模时间序列数据的已知和隐藏的非线性特征。本文报告了一个案例研究,其中选择的金融和股票的时间序列数据的70多个股票指数聚类成不同的群体,使用介绍的两阶段程序。结果表明,该方法能够实现87.5%的准确率聚类和预测标签看不见的时间序列数据。该论文还报告了一个重要的发现,其中观察到这两种技术的性能(即,autoencoder和Kmeans)具有可比性。然而,与Kmeans算法相比,基于自动编码器的方法对一些时间序列数据进行了不同的分类。结果可能表明,拟议的基于深度学习的方法正在考虑传统Kmeans可能会忽视的其他隐藏特征。这一发现提出了一个问题,即是否应该分析数据的显式特征以进行聚类,或者是否需要采用更先进的技术(如深度学习)来探索隐藏的特征和关系以进行聚类。
This paper introduces a two-stage deep learning-based methodology for clustering time series data. First, a novel technique is introduced to utilize the characteristics (e.g., volatility) of the given time series data in order to create labels and thus enable transformation of the problem from an unsupervised into a supervised learning. Second, an autoencoder-based deep learning model is built to model both known and hidden non-linear features of time series data. The paper reports a case study in which the selected financial and stock time series data of over 70 stock indices are clustered into distinct groups using the introduced two-stage procedure. The results show that the proposed methodology is capable of achieving 87.5% accuracy in clustering and predicting the labels for unseen time series data. The paper also reports an important finding in which it is observed that the performance of both techniques (i.e., autoencoder and Kmeans) are comparable. However, there are a few instances of time series data that are classified differently by the autoencoder-based methodology compared to the Kmeans algorithm. The results may indicate that the proposed deep learning-based approach is taking into account additional hidden features that might be overlooked by conventional Kmeans. The finding raises the question whether the explicit features of data should be analyzed for clustering or more advanced techniques such as deep learning need to be adapted by which hidden features and relationships are explored for clustering purposes.