Naive Bayes for Text Classification with Unbalanced Classes

Naive Bayes for Text Classification with Unbalanced Classes
复制标题

DOI:
10.1007/11871637_49
复制
发表时间:
2006-09
期刊:
--
影响因子:
--
通讯作者:
E. Frank;R. Bouckaert
E. Frank;R. Bouckaert
中科院分区:
其他
文献类型:
--
作者:
E. Frank;R. Bouckaert

文献摘要

被引文献

相似文献

多项朴素贝叶斯(MNB)是一种流行的文档分类方法,由于其计算效率和相对较好的预测性能。最近已经确定,预测性能可以通过适当的数据转换进一步提高[1,2]。在本文中,我们提出了另一种转换,旨在解决将MNB应用于不平衡数据集的潜在问题。我们提出了一个适当的修正,通过调整属性先验。这种校正可以作为另一个数据标准化步骤来实现,我们表明它可以显着改善ROC曲线下面积。我们还表明,修改后的版本的MNB是非常密切相关的简单的基于质心的分类器,并比较这两种方法的经验。
Multinomial naive Bayes (MNB) is a popular method for document classification due to its computational efficiency and relatively good predictive performance. It has recently been established that predictive performance can be improved further by appropriate data transformations [1,2]. In this paper we present another transformation that is designed to combat a potential problem with the application of MNB to unbalanced datasets. We propose an appropriate correction by adjusting attribute priors. This correction can be implemented as another data normalization step, and we show that it can significantly improve the area under the ROC curve. We also show that the modified version of MNB is very closely related to the simple centroid-based classifier and compare the two methods empirically.