Feature selection for text classification with Naive Bayes

Feature selection for text classification with Naive Bayes
复制标题

使用朴素贝叶斯进行文本分类的特征选择

DOI:
10.1016/j.eswa.2008.06.054
复制
发表时间:
2009-04-01
影响因子:
8.5
通讯作者:
Qu, Youli
Qu, Youli
中科院分区:
计算机科学1区
文献类型:
--
作者:
Chen, Jingnian;Huang, Houkuan;Qu, Youli

文献摘要

被引文献

相似文献

特征选择作为文本分类中一项重要的预处理技术,可以提高文本分类器的可扩展性、效率和准确率。一般来说,一个好的特征选择方法应该考虑域和算法特性。由于朴素贝叶斯分类器简单、高效,且对特征选择高度敏感,因此对其进行特征选择的研究具有重要意义。本文提出了两个用于多类文本数据集的朴素贝叶斯分类器的特征评价指标:多类比值比(莫尔)和类区分度(CDM)。在两个多类文本集上进行了朴素贝叶斯分类器的文本分类实验。实验结果表明,CDM和莫尔的选择效果明显优于其他特征选择方法。(C)2008爱思唯尔有限公司保留所有权利。
As an important preprocessing technology in text classification, feature selection can improve the scalability, efficiency and accuracy of a text classifier. In general, a good feature selection method should consider domain and algorithm characteristics. As the Naive Bayesian classifier is very simple and efficient and highly sensitive to feature selection, so the research of feature selection specially for it is significant. This paper presents two feature evaluation metrics for the Naive Bayesian classifier applied on multi-class text datasets: Multi-class Odds Ratio (MOR), and Class Discriminating Measure (CDM). Experiments of text classification with Naive Bayesian classifiers were carried out on two multi-class texts collections. As the results indicate, CDM and MOR gain obviously better selecting effect than other feature selection approaches. (C) 2008 Elsevier Ltd. All rights reserved.