Relevance popularity: A term event model based feature selection scheme for text classification.

Relevance popularity: A term event model based feature selection scheme for text classification.
复制标题

相关流行度:基于术语事件模型的文本分类特征选择方案

DOI:
10.1371/journal.pone.0174341
复制
发表时间:
2017
期刊:
影响因子:
3.7
通讯作者:
Zhang L
Zhang L
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Feng G;An B;Yang F;Wang H;Zhang L

文献摘要

被引文献

相似文献

特征选择是通过优化输入到分类器的特征子集来提高文本分类方法性能的一种实用方法。在信息增益和卡方等传统的特征选择方法中,经常使用包含特定术语的文档数量(即文档频率)。然而,一个特定术语在每一份文件中出现的频率尚未得到充分调查,尽管这是一个很有希望的特征,可以产生准确的分类。本文提出了一种新的基于项事件多项式朴素贝叶斯概率模型的特征选择方法。根据模型假设,基于预测概率比的匹配得分函数可以分解。最后,我们用每一项的估计量替换其内部参数,得到每一项的特征选择度量。在一个基准英文文本集(20个新闻组)和一个中文文本集(MPH-20)上,我们使用两个广泛使用的文本分类器(朴素贝叶斯和支持向量机)进行了数值实验,结果表明,我们的方法比典型的特征选择方法具有更好的性能。
Feature selection is a practical approach for improving the performance of text classification methods by optimizing the feature subsets input to classifiers. In traditional feature selection methods such as information gain and chi-square, the number of documents that contain a particular term (i.e. the document frequency) is often used. However, the frequency of a given term appearing in each document has not been fully investigated, even though it is a promising feature to produce accurate classifications. In this paper, we propose a new feature selection scheme based on a term event Multinomial naive Bayes probabilistic model. According to the model assumptions, the matching score function, which is based on the prediction probability ratio, can be factorized. Finally, we derive a feature selection measurement for each term after replacing inner parameters by their estimators. On a benchmark English text datasets (20 Newsgroups) and a Chinese text dataset (MPH-20), our numerical experiment results obtained from using two widely used text classifiers (naive Bayes and support vector machine) demonstrate that our method outperformed the representative feature selection methods.