A New Feature Selection Score for Multinomial Naive Bayes Text Classification Based on KL-Divergence

A New Feature Selection Score for Multinomial Naive Bayes Text Classification Based on KL-Divergence
复制标题

DOI:
10.3115/1219044.1219068
复制
发表时间:
2004-07
期刊:
--
影响因子:
--
通讯作者:
Karl-Michael Schneider
Karl-Michael Schneider
中科院分区:
其他
文献类型:
--
作者:
Karl-Michael Schneider

文献摘要

被引文献

相似文献

我们基于训练文档中单词分布与其类别之间的kl散度定义了一个新的文本分类特征选择分数。分数倾向于在同一类文档中具有相似分布但在不同类文档中具有不同分布的单词。在两个标准数据集上的实验表明,新方法优于互信息,特别是对于较小的类别。
We define a new feature selection score for text classification based on the KL-divergence between the distribution of words in training documents and their classes. The score favors words that have a similar distribution in documents of the same class but different distributions in documents of different classes. Experiments on two standard data sets indicate that the new method outperforms mutual information, especially for smaller categories.