Information Theory Inspired Pattern Analysis for Time-series Data

Information Theory Inspired Pattern Analysis for Time-series Data
复制标题

DOI:
10.48550/arxiv.2302.11654
复制
发表时间:
2023-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Yushan Huang;Yuchen Zhao;Alexander Capstick;Francesca Palermo;Hamed Haddadi;P. Barnaghi
Yushan Huang;Yuchen Zhao;Alexander Capstick;Francesca Palermo;Hamed Haddadi;P. Barnaghi
中科院分区:
其他
文献类型:
--
作者:
Yushan Huang;Yuchen Zhao;Alexander Capstick;Francesca Palermo;Hamed Haddadi;P. Barnaghi

文献摘要

相似文献

目前的时间序列模式分析方法主要依靠统计特征或概率学习和推理方法来识别数据中的模式和趋势。当应用于多变量、多源、状态变化和有噪声的时间序列数据时,这些方法不能很好地推广。为了解决这些问题,我们提出了一种高度泛化的方法,该方法使用基于信息论的特征来识别和学习多变量时间序列数据中的模式。为了演示所提出的方法,我们分析了人类活动数据中的模式变化。对于具有随机状态转移的应用,基于马尔可夫链的Shannon熵、马尔可夫链的熵率、马尔可夫链的熵积和马尔可夫链的von Neumann熵来开发特征。对于状态建模不适用的应用,我们利用了五种熵变量,包括近似熵、增量熵、离散度熵、相位熵和斜率熵。结果表明,基于信息论的特征比基线模型的召回率、F1评分和准确率平均提高了23.01%,模型结构更简单,模型参数的个数平均减少了18.75倍。
Current methods for pattern analysis in time series mainly rely on statistical features or probabilistic learning and inference methods to identify patterns and trends in the data. Such methods do not generalize well when applied to multivariate, multi-source, state-varying, and noisy time-series data. To address these issues, we propose a highly generalizable method that uses information theory-based features to identify and learn from patterns in multivariate time-series data. To demonstrate the proposed approach, we analyze pattern changes in human activity data. For applications with stochastic state transitions, features are developed based on Shannon's entropy of Markov chains, entropy rates of Markov chains, entropy production of Markov chains, and von Neumann entropy of Markov chains. For applications where state modeling is not applicable, we utilize five entropy variants, including approximate entropy, increment entropy, dispersion entropy, phase entropy, and slope entropy. The results show the proposed information theory-based features improve the recall rate, F1 score, and accuracy on average by up to 23.01% compared with the baseline models and a simpler model structure, with an average reduction of 18.75 times in the number of model parameters.