On time series classification with dictionary-based classifiers

On time series classification with dictionary-based classifiers
复制标题

DOI:
10.3233/ida-184333
复制
发表时间:
2019-10
期刊:
Intell. Data Anal.
影响因子:
--
通讯作者:
J. Large;A. Bagnall;S. Malinowski;R. Tavenard
J. Large;A. Bagnall;S. Malinowski;R. Tavenard
中科院分区:
其他
文献类型:
--
作者:
J. Large;A. Bagnall;S. Malinowski;R. Tavenard

文献摘要

被引文献

相似文献

用于时间序列分类(TSC)的一系列算法涉及在每个序列上运行滑动窗口,将窗口离散化以形成单词,在词典上形成单词计数的直方图,然后在直方图上构建分类器。最近对两种这类算法,模式袋(BOP)和符号傅里叶近似符号袋(BOSS)的评估发现,这些看似相似的算法在精度上存在显著差异。我们通过解构分类器并测量BOSS和BOSS之间四个关键组成部分的相对重要性来研究这一现象。我们发现,虽然集成是这两种算法的关键组成部分,但其他组成部分的影响是混合的,而且更复杂。我们得出的结论是,BOSS代表了基于词典的TSC的最新技术。BOSS和BOSS都可以归类为词袋逼近。这在计算机视觉中特别流行,用于图像分类等任务。我们将计算机视觉中使用的三种技术应用于TSC:尺度不变特征变换、空间金字塔和直方图相交。我们发现,将空间金字塔与BOSS(SP)结合使用可以产生明显更准确的分类器。SP比标准基准和原始BOSS算法的精确度要高得多。它并不比最好的基于形状的学习方法或深度学习方法差多少,只有在将BOSS作为组成模块的集成中表现更好。
A family of algorithms for time series classification (TSC) involve running a sliding window across each series, discretising the window to form a word, forming a histogram of word counts over the dictionary, then constructing a classifier on the histograms. A recent evaluation of two of this type of algorithm, Bag of Patterns (BOP) and Bag of Symbolic Fourier Approximation Symbols (BOSS) found a significant difference in accuracy between these seemingly similar algorithms. We investigate this phenomenon by deconstructing the classifiers and measuring the relative importance of the four key components between BOP and BOSS. We find that whilst ensembling is a key component for both algorithms, the effect of the other components is mixed and more complex. We conclude that BOSS represents the state of the art for dictionary-based TSC. Both BOP and BOSS can be classed as bag of words approaches. These are particularly popular in Computer Vision for tasks such as image classification. We adapt three techniques used in Computer Vision for TSC: Scale Invariant Feature Transform; Spatial Pyramids; and Histogram Intersection. We find that using Spatial Pyramids in conjunction with BOSS (SP) produces a significantly more accurate classifier. SP is significantly more accurate than standard benchmarks and the original BOSS algorithm. It is not significantly worse than the best shapelet-based or deep learning approaches, and is only outperformed by an ensemble that includes BOSS as a constituent module.