Feature ranking for multi-label classification using Markov networks

Feature ranking for multi-label classification using Markov networks
复制标题

DOI:
10.1016/j.neucom.2016.04.023
复制
发表时间:
2016-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Paweł Teisseyre
Paweł Teisseyre
中科院分区:
其他
文献类型:
--
作者:
Paweł Teisseyre

文献摘要

被引文献

相似文献

提出了一种简单有效的多标签分类特征排序方法。该方法产生一个特征排名,显示它们在预测标签时的相关性,这反过来又允许我们选择最终的特征子集。该过程基于马尔可夫网络,允许我们以直接的方式对标签和特征之间的依赖关系进行建模。在第一步中,我们只使用标签构建一个简单的网络,然后测试添加单个特征对初始网络的影响程度。更具体地说,在第一步中,我们使用Ising模型,而第二步是基于分数统计,这允许我们非常快速地测试添加功能的重要性。该方法不需要对标签空间进行变换,给出了可解释的结果,并允许有吸引力的依赖结构可视化。我们通过讨论伊辛模型和分数统计量的一些理论性质,给出了这一过程的理论证明。文中还讨论了基于Ising模型的正则化Logistic回归特征排序方法。数值实验表明,该方法在人工数据集和真实数据集上的性能均优于传统方法。
We propose a simple and efficient method for ranking features in multi-label classification. The method produces a ranking of features showing their relevance in predicting labels, which in turn allows us to choose a final subset of features. The procedure is based on Markov networks and allows us to model the dependencies between labels and features in a direct way. In the first step we build a simple network using only labels and then we test how much adding a single feature affects the initial network. More specifically, in the first step we use the Ising model whereas the second step is based on the score statistic, which allows us to test a significance of added features very quickly. The proposed approach does not require transformation of label space, gives interpretable results and allows for attractive visualization of dependency structure. We give a theoretical justification of the procedure by discussing some theoretical properties of the Ising model and the score statistic. We also discuss feature ranking procedure based on fitting Ising model usingl1regularized logistic regressions. Numerical experiments show that the proposed methods outperform the conventional approaches on the considered artificial and real datasets.