Feature Selection with the R Package MXM: Discovering Statistically Equivalent Feature Subsets

Feature Selection with the R Package MXM: Discovering Statistically Equivalent Feature Subsets
复制标题

DOI:
10.18637/jss.v080.i07
复制
发表时间:
2017-09-01
影响因子:
5.8
通讯作者:
Tsamardinos, Ioannis
Tsamardinos, Ioannis
中科院分区:
计算机科学2区
文献类型:
--
作者:
Lagani, Vincenzo;Athineou, Giorgos;Tsamardinos, Ioannis

文献摘要

被引文献

相似文献

统计等价签名(SES)算法是一种受贝叶斯网络基于约束学习原理启发的特征选择方法。目前可用的大多数特征选择方法只返回一个特征子集,据说是具有最高预测能力的一个。我们认为,在几个领域的多个子集可以实现接近最大的预测精度,任意提供只有一个有几个缺点。SES方法试图识别多个预测特征子集,其性能在统计上是等效的。在这方面,SES算法包含并扩展了以前的特征选择算法,如最大-最小父子算法。SES算法是在R包MXM中包含的同名函数中实现的,代表mens ex machina,在拉丁语中的意思是“来自机器的思想”。SES的MXM实现处理多个数据分析任务,即分类、回归和生存分析。在本文中,我们介绍了SES算法,其实现,并提供了使用的SES功能在R。此外,我们分析了三个公开的数据集来说明SES检索的签名的等效性,并将SES与最先进的特征选择方法LASSO进行对比。我们的研究结果提供了初步的证据表明,这两种方法在预测准确性方面表现良好,并且在真实的世界数据中实际存在多个同样具有预测性的签名。
The statistically equivalent signature (SES) algorithm is a method for feature selection inspired by the principles of constraint-based learning of Bayesian networks. Most of the currently available feature selection methods return only a single subset of features, supposedly the one with the highest predictive power. We argue that in several domains multiple subsets can achieve close to maximal predictive accuracy, and that arbitrarily providing only one has several drawbacks. The SES method attempts to identify multiple, predictive feature subsets whose performances are statistically equivalent. In that respect the SES algorithm subsumes and extends previous feature selection algorithms, like the max-min parent children algorithm.The SES algorithm is implemented in an homonym function included in the R package MXM, standing for mens ex machina, meaning 'mind from the machine' in Latin. The MXM implementation of SES handles several data analysis tasks, namely classification, regression and survival analysis. In this paper we present the SES algorithm, its implementation, and provide examples of use of the SES function in R. Furthermore, we analyze three publicly available data sets to illustrate the equivalence of the signatures retrieved by SES and to contrast SES against the state-of-the-art feature selection method LASSO. Our results provide initial evidence that the two methods perform comparably well in terms of predictive accuracy and that multiple, equally predictive signatures are actually present in real world data.