A Bayesian feature selection paradigm for text classification

A Bayesian feature selection paradigm for text classification
复制标题

用于文本分类的贝叶斯特征选择范例

DOI:
10.1016/j.ipm.2011.08.002
复制
发表时间:
2012-03-01
影响因子:
8.6
通讯作者:
Hao, Lizhu
Hao, Lizhu
中科院分区:
计算机科学1区
文献类型:
--
作者:
Feng, Guozhong;Guo, Jianhua;Hao, Lizhu

文献摘要

被引文献

相似文献

由于数字形式文档的可用性不断增加以及随之而来的组织文档的需求,将文本自动分类为预定义类别引起了人们的广泛关注。文本分类的一个重要问题是特征选择,其目标是提高分类效果和/或计算效率。由于社交文本集合中分类不平衡和特征稀疏,过滤方法可能效果不佳。在本文中,我们在训练过程中进行特征选择,通过从一组预先分类的文档中学习类别的特征来自动选择最佳特征子集。我们提出了一种生成概率模型,通过分布描述类别,通过引入二元排除/包含潜在向量来处理特征选择问题,该向量通过高效的 Metropolis 搜索进行更新。现实生活中的例子说明了该方法的有效性。 (C) 2011 Elsevier Ltd. 保留所有权利。
The automated classification of texts into predefined categories has witnessed a booming interest, due to the increased availability of documents in digital form and the ensuing need to organize them. An important problem for text classification is feature selection, whose goals are to improve classification effectiveness, computational efficiency, or both. Due to categorization unbalancedness and feature sparsity in social text collection, filter methods may work poorly. In this paper, we perform feature selection in the training process, automatically selecting the best feature subset by learning, from a set of preclassified documents, the characteristics of the categories. We propose a generative probabilistic model, describing categories by distributions, handling the feature selection problem by introducing a binary exclusion/inclusion latent vector, which is updated via an efficient Metropolis search. Real-life examples illustrate the effectiveness of the approach. (C) 2011 Elsevier Ltd. All rights reserved.