A rough set-based hybrid method to text categorization

A rough set-based hybrid method to text categorization
复制标题

一种基于粗糙集的文本分类混合方法

DOI:
10.1109/wise.2001.996486
复制
发表时间:
2001
期刊:
Proceedings of the Second International Conference on Web Information Systems Engineering
影响因子:
--
通讯作者:
N. Ishii
N. Ishii
中科院分区:
--
文献类型:
--
作者:
Y. Bao;Satoshi Aoyama;Xiaoyong Du;Kazutaka Yamada;N. Ishii

文献摘要

参考文献

被引文献

相似文献

本文提出了一种基于粗糙集理论的混合文本分类方法。在信息过滤和检索(IF/IR)的良好的文本分类的一个中心问题是高维数据。它可能包含许多不必要和不相关的功能。为了科普这个问题,我们提出了一种混合技术,使用潜在语义索引(LSI)和粗糙集理论(RS),以缓解这种情况。给定文档的语料库和分类文档的训练集的示例,该技术定位坐标关键字的最小集合以区分文档的类别,从而降低关键字向量的维度。这简化了基于知识的IF/IR系统的创建,加快了它们的操作,并允许对所采用的规则库进行轻松编辑。此外,我们产生多个知识库,而不是一个知识库的分类新的对象,希望多个知识库的答案的组合导致更好的性能。多个知识库可以在RS的框架内以统一的方式精确地表达。本文描述了所提出的技术,讨论了集成的关键字获取算法,潜在语义索引(LSI)与基于粗糙集的规则生成算法,并提供了实验结果。实验结果表明,该方法优于粗糙集方法。
In this paper we present a hybrid text categorization method based on Rough Sets theory. A central problem in good text classification for information filtering and retrieval (IF/IR) is the high dimensionality of the data. It may contain many unnecessary and irrelevant features. To cope with this problem, we propose a hybrid technique using Latent Semantic Indexing (LSI) and Rough Sets theory (RS) to alleviate this situation. Given corpora of documents and a training set of examples of classified documents, the technique locates a minimal set of co-ordinate keywords to distinguish between classes of documents, reducing the dimensionality of the keyword vectors. This simplifies the creation of knowledge-based IF/IR systems, speeds up their operation, and allows easy editing of the rule bases employed. Besides, we generate several knowledge base instead of one knowledge base for the classification of new object, hoping that the combination of answers of the multiple knowledge bases result in better performance. Multiple knowledge bases can be formulated precisely and in a unified way within the framework of RS. This paper describes the proposed technique, discusses the integration of a keyword acquisition algorithm, Latent Semantic indexing (LSI) with Rough Set-based rule generate algorithm, and provides experimental results. The test results show the hybrid method is better than the previous rough set-based approach.
DOI: 10.1080/01638539809545028
发表时间: 1998-01-01
影响因子: 2.2
作者:
Landauer, TK;Foltz, PW;Laham, D
通讯作者: Laham, D