Feature screening for ultrahigh dimensional binary data

Feature screening for ultrahigh dimensional binary data
复制标题

超高维二进制数据的特征筛选

DOI:
10.4310/sii.2018.v11.n1.a4
复制
发表时间:
2018-01-01
影响因子:
0.8
通讯作者:
Guo, Jianhua
Guo, Jianhua
中科院分区:
数学4区
文献类型:
--
作者:
Guan, Guoyu;Shan, Na;Guo, Jianhua

文献摘要

被引文献

相似文献

随着信息技术的飞速发展,二维二值数据急剧增加,特征筛选已成为真实的数据分析中必不可少的一步。在这篇文章中,我们提出了一个L-0-正则化的朴素贝叶斯分类器的特征筛选过程,这是等价于经典的互信息筛选方法。然而,L-0正则化中的转向参数选择困难,缺乏理论支持。为此,BIC型标准被应用于识别重要特征。此外,在一些温和的假设下,所提出的方法的渐近性质进行了理论研究。最后,在模拟数据上验证了该算法的优良性能,并给出了一个真实的中文文档分类实例。
With the rapid development of information technology, ultrahigh dimensional binary data have increased dramatically, for which feature screening has become a necessary step in real data analysis. In this article, we propose a L-0-regularization feature screening procedure for naive Bayes classifier, which is equivalent to the classical mutual information screening method. However, the turning parameter in L-0-regularization is hard to be selected and lack of theoretical support. To this end, a BIC-type criterion is applied to identify important features. Moreover, the asymptotic properties of the proposed method is theoretically investigated under some mild assumptions. Lastly, its outstanding performance is numerically confirmed on simulated data, and a real example of Chinese document classification is presented for illustration purpose.