ASFS: A novel streaming feature selection for multi-label data based on neighborhood rough set

ASFS: A novel streaming feature selection for multi-label data based on neighborhood rough set
复制标题

DOI:
10.1007/s10489-022-03366-x
复制
发表时间:
2022-05-02
影响因子:
5.3
通讯作者:
Zhang, Jia
Zhang, Jia
中科院分区:
计算机科学2区
文献类型:
--
作者:
Liu, Jinghua;Lin, Yaojin;Zhang, Jia

文献摘要

被引文献

相似文献

基于邻域粗糙集的在线流特征选择方法近年来受到广泛关注,在高维数据处理中发挥了重要作用。然而,现有的方法大多直接应用于处理单标签数据,或者通过将多标签数据转换为多个单标签数据集的组合来处理多标签数据,忽略了多标签数据的标签集是一个完整的整体。在本文中,我们提出了一种新的在线流特征选择的多标签学习通过neighborhoorough集模型,其中特征的重要性,特征冗余和标签空间的完整性,同时考虑。具体而言,首先定义了一种新的自适应邻域关系,避免了邻域参数的设置,并将邻域粗糙集模型重构为适合直接处理多标签数据的模型。在此基础上,引入一个评价准则来选择相对于标签集和当前选择的特征重要的特征,并提出一个优化目标函数来更新选择的特征子集和过滤掉冗余特征。在不同类型数据集上的对比实验明确地验证了所提出的方法的优点。
Neighborhood rough set based online streaming feature selection methods have aroused wide concern in recent years and played a vital role in processing high-dimensional data. However, most of the existing methods are directly applied to handle single-label data, or to handle multi-label data by converting multi-label data into a combination of multiple single-label datasets, which ignores that the label set of multi-label data is an integral whole. In this paper, we propose a novel online streaming feature selection for multi-label learning via the neighborhoorough set model, in which feature significance, feature redundancy, and label space integrity are taken into account, simultaneously. To be specific, we first define a new adaptive neighborhood relation to avoid the setting of neighborhood parameter and restructure the neighborhood rough set model to be suitable for processing multi-label data directly. Based on this model, we introduce a evaluation criterion to select features that are important relative to label set and the currently selected features, and present an optimization objective function to update the selected feature subset and filter out redundant features. Comparative experiments on different types of data sets explicitly verify the advantages of the proposed method.