Incremental Learning in Time-series Data using Reinforcement Learning

Incremental Learning in Time-series Data using Reinforcement Learning
复制标题

DOI:
10.1109/icdmw58026.2022.00115
复制
发表时间:
2022-11
期刊:
2022 IEEE International Conference on Data Mining Workshops (ICDMW)
影响因子:
--
通讯作者:
Mustafa Shuqair;J. Jimenez-shahed;Behnaz Ghoraani
Mustafa Shuqair;J. Jimenez-shahed;Behnaz Ghoraani
中科院分区:
其他
文献类型:
--
作者:
Mustafa Shuqair;J. Jimenez-shahed;Behnaz Ghoraani

文献摘要

相似文献

随着可穿戴传感器和连续监控工具的不断增长,系统监控已成为人们关注的领域。然而,分类模型对未见过的输入数据的通用性仍然具有挑战性。本文提出了一种基于强化学习(RL)的新颖架构,以增量学习时间序列数据的模式并检测系统状态的变化。我们的理由是,强化学习从过去的经验中学习的能力可以帮助提高时间序列监控应用中分类模型的性能和通用性。我们对环境的新颖定义包括一组一类异常检测器,用于根据传入数据的动态定义环境状态,以及一个奖励函数,用于根据 RL 代理的行为对其进行奖励。深度强化学习代理逐步学习根据环境状态和收到的奖励执行连续的二元分类预测。我们应用所提出的模型来检测帕金森病 (PD) 患者对药物(开或关)的反应。 PD 数据集包含使用两个可穿戴传感器从 12 名患者收集的 170 分钟时间序列运动信号。我们提出的模型的测试精度为 77.95%,优于自适应增强、多层感知器和支持向量机,测试精度分别为 53.10%、44.92% 和 52.70%。所提出的模型的 F 分数略有下降,从验证分数的 88.15% 下降到测试中的 78.42%,与其他三个模型相比,下降幅度明显较小。这些证明了所提出的基于强化学习的分类器在时间序列监控应用中作为未见输入数据的高度通用模型的潜力。
System monitoring has become an area of interest with the increasing growth in wearable sensors and continuous monitoring tools. However, the generalizability of the classification models to unseen incoming data remains challenging. This paper proposes a novel architecture based on reinforcement learning (RL) to incre-mentally learn patterns of time-series data and detect changes in the system state. Our rationale is that RL's ability to learn from past experiences can help increase the performance and generalizability of classification models in time-series monitoring applications. Our novel definition of the environment consists of a set of one-class anomaly detectors to define environment states based on the dynamics of the incoming data and a reward function to reward the RL agent according to its actions. A deep RL agent incrementally learns to perform continuous, binary classification predictions according to the environment states and the received reward. We applied the proposed model for detecting response to medication (ON or OFF) in patients with Parkinson's disease (PD). The PD dataset consisted of 170 minutes of time-series movement signals collected from 12 patients using two wearable sensors. Our proposed model, with a testing accuracy of 77.95%, outperformed Adaptive Boosting, Multi-layer Perceptron, and Support Vector Machines with 53.10%, 44.92%, and 52.70% testing accuracy, respectively. The proposed model had a slight decline in the F-score, decreasing from 88.15% validation score to 78.42% in testing, a significantly slight decline compared to the other three models. These evidence the potential of the proposed RL-based classifier in time-series monitoring applications as a highly generalizable model for unseen incoming data.