Feature selection, perceptron learning, and a usability case study for text categorization

Feature selection, perceptron learning, and a usability case study for text categorization
复制标题

DOI:
10.1145/258525.258537
复制
发表时间:
1997-07
期刊:
--
影响因子:
--
通讯作者:
H. Ng;Wei Boon Goh;Kok Leong Low
H. Ng;Wei Boon Goh;Kok Leong Low
中科院分区:
其他
文献类型:
--
作者:
H. Ng;Wei Boon Goh;Kok Leong Low

文献摘要

被引文献

相似文献

在本文中,我们描述了一个自动学习方法的文本分类感知学习和一个新的特征选择度量,称为相关系数的基础上。我们的方法已经在标准的路透社文本分类集合上进行了测试。实证结果表明,我们的方法优于此% uters集合上发布的最佳结果。特别是,我们的新的特征选择方法产生comiderable改进。我们还调查了我们的自动化hxu-n-~方法的可用性,实际上开发了一个系统,分类文本到一个树的类别。我们比较了我们的学习方法的准确性,一个rrdmdmsed的,专家系统ap preach,使用由Cams gie Group构建的文本分类外壳。虽然我们的自动学习方法仍然给出了较低的准确性,但通过适当地引入一组手动选择的词来用作fure,组合的半自动方法产生接近于 * baaed方法的准确性。
In this paper, we describe an automated learning approach to text categorization based on perception learning and a new feature selection metric, called correlation coefficient. Our approach has been teated on the standard Reuters text categorization collection. Empirical results indicate that our approach outperforms the best published results on this % uters collection. In particular, our new feature selection method yields comiderable improvement. We also investigate the usability of our automated hxu-n-~ approach by actually developing a system that categorizes texts into a tree of categories. We compare tbe accuracy of our learning approach to a rrddmsed, expert system ap preach that uses a text categorization shell built by Cams gie Group. Although our automated learning approach still gives a lower accuracy, by appropriately inmrporating a set of manually chosen worda to use as f~ures, the combined, semi-automated approach yields accuracy close to the * baaed approach.