Interactive reinforced feature selection with traverse strategy

Interactive reinforced feature selection with traverse strategy
复制标题

DOI:
10.1007/s10115-022-01812-3
复制
发表时间:
2023-01-21
影响因子:
2.7
通讯作者:
Fu,Yanjie
Fu,Yanjie
中科院分区:
计算机科学4区
文献类型:
--
作者:
Liu,Kunpeng;Wang,Dongjie;Fu,Yanjie

文献摘要

相似文献

在本文中,我们提出了一种基于单智能体蒙特卡罗的增强特征选择方法,以及两个效率提高策略,即,提前停止策略和奖励级交互策略。特征选择是数据预处理中最重要的技术之一,旨在为给定的下游机器学习任务找到最优的特征子集。为了提高其效力和效率,进行了大量研究。近年来,多智能体增强特征选择(MARFS)在提高特征选择性能方面取得了巨大的成功。然而,MARFS的计算成本负担很重,这大大限制了它在现实世界中的应用。在本文中,我们提出了一种有效的增强特征选择方法,它使用一个代理遍历整个特征集,并决定选择或不选择每个特征逐一。具体来说,我们首先开发一个行为策略,并使用它来遍历特征集并生成训练数据。然后,根据训练数据对目标策略进行评估,并利用Bellman方程对目标策略进行改进。此外,我们以增量的方式进行重要性采样,并提出了一个早期停止策略,以提高训练效率,通过消除偏斜数据。在早期停止策略中,行为策略停止遍历的概率与重要性采样权重成反比。此外,我们提出了一个奖励层面和培训层面的互动策略,以提高培训效率,通过外部的建议。此外,我们提出了一种增量描述统计方法来表示状态,具有较低的计算成本。最后,我们设计了大量的实验,在现实世界的数据来证明所提出的方法的优越性。
In this paper, we propose a single-agent Monte Carlo-based reinforced feature selection method, as well as two efficiency improvement strategies, i.e., early stopping strategy and reward-level interactive strategy. Feature selection is one of the most important technologies in data prepossessing, aiming to find the optimal feature subset for a given downstream machine learning task. Enormous research has been done to improve its effectiveness and efficiency. Recently, the multi-agent reinforced feature selection (MARFS) has achieved great success in improving the performance of feature selection. However, MARFS suffers from the heavy burden of computational cost, which greatly limits its application in real-world scenarios. In this paper, we propose an efficient reinforcement feature selection method, which uses one agent to traverse the whole feature set and decides to select or not select each feature one by one. Specifically, we first develop one behavior policy and use it to traverse the feature set and generate training data. And then, we evaluate the target policy based on the training data and improve the target policy by Bellman equation. Besides, we conduct the importance sampling in an incremental way and propose an early stopping strategy to improve the training efficiency by the removal of skew data. In the early stopping strategy, the behavior policy stops traversing with a probability inversely proportional to the importance sampling weight. In addition, we propose a reward-level and training-level interactive strategy to improve the training efficiency via external advice. What’s more, we propose an incremental descriptive statistics method to represent the state with low computational cost. Finally, we design extensive experiments on real-world data to demonstrate the superiority of the proposed method.