Highly Comparative Feature-Based Time-Series Classification

Highly Comparative Feature-Based Time-Series Classification
复制标题

DOI:
10.1109/tkde.2014.2316504
复制
发表时间:
2014-12-01
影响因子:
8.9
通讯作者:
Jones, Nick S.
Jones, Nick S.
中科院分区:
计算机科学2区
文献类型:
--
作者:
Fulcher, Ben D.;Jones, Nick S.

文献摘要

被引文献

相似文献

介绍了一种基于特征的高度比较性的时间序列分类方法,该方法使用一个广泛的算法数据库从时间序列中提取数千个可解释的特征。这些特征源自科学时间序列分析文献,包括时间序列在相关性结构、分布、熵、平稳性、缩放特性以及对一系列时间序列模型的拟合等方面的汇总。在为训练集中的每个时间序列计算数千个特征之后,使用带有线性分类器的贪心前向特征选择来挑选那些对类别结构最具信息量的特征。由此产生的基于特征的分类器使用数量减少的时间序列特性自动学习类别之间的差异,并且避免了计算时间序列之间距离的需要。以这种方式表示时间序列导致维度大幅降低,使得该方法在包含长时间序列或不同长度时间序列的非常大的数据集上表现良好。对于所研究的许多数据集,分类性能超过了传统的基于实例的分类器,包括使用欧几里得距离和动态时间规整的最近邻分类器,并且最重要的是,所选特征提供了对数据集特性的理解,这种洞察力可以指导进一步的科学研究。
A highly comparative, feature-based approach to time series classification is introduced that uses an extensive database of algorithms to extract thousands of interpretable features from time series. These features are derived from across the scientific time-series analysis literature, and include summaries of time series in terms of their correlation structure, distribution, entropy, stationarity, scaling properties, and fits to a range of time-series models. After computing thousands of features for each time series in a training set, those that are most informative of the class structure are selected using greedy forward feature selection with a linear classifier. The resulting feature-based classifiers automatically learn the differences between classes using a reduced number of time-series properties, and circumvent the need to calculate distances between time series. Representing time series in this way results in orders of magnitude of dimensionality reduction, allowing the method to perform well on very large data sets containing long time series or time series of different lengths. For many of the data sets studied, classification performance exceeded that of conventional instance-based classifiers, including one nearest neighbor classifiers using euclidean distances and dynamic time warping and, most importantly, the features selected provide an understanding of the properties of the data set, insight that can guide further scientific investigation.