Text Classification: Combining Grouping, LSA and kNN vs Support Vector Machine

Text Classification: Combining Grouping, LSA and kNN vs Support Vector Machine
复制标题

DOI:
10.1007/11893004_51
复制
发表时间:
2006-10
期刊:
--
影响因子:
--
通讯作者:
N. Ishii;Takeshi Murai;T. Yamada;Y. Bao;Susumu Suzuki
N. Ishii;Takeshi Murai;T. Yamada;Y. Bao;Susumu Suzuki
中科院分区:
其他
文献类型:
--
作者:
N. Ishii;Takeshi Murai;T. Yamada;Y. Bao;Susumu Suzuki

文献摘要

被引文献

相似文献

文本分类是处理和组织文本数据的关键技术。支持向量机(SVM)被证明是更好的知名方法之间的分类。本文提出了一种相似词的分组方法,并将其应用于路透社新闻的分类,结果表明,该方法在分类精度上与潜在语义分析(LSA)方法相当。此外,提出了一种新的组合方法的分类,它包括的聚类分析,LSA其次是k-最近邻分类(k-NN)。这里提出的组合方法,显示出更高的精度比传统的方法的kNN,和LSA其次是kNN的分类。然后,组合方法显示出几乎相同的精度支持向量机。
Text classification is a key technique for handling and organizing text data. The support vector machine(SVM) is shown to be better for the classification among well-known methods. In this paper, the grouping method of the similar words, is proposed for the classification of documents, which is applied to Reuters news and it is shown that the grouping of words has equivalent ability to the Latent Semantic Analysis(LSA) in the classification accuracy. Further, a new combining method is proposed for the classification, which consists of Grouping, LSA followed by the k-Nearest Neighbor classification ( k-NN ). The combining method proposed here, shows the higher accuracy in the classification than the conventional methods of the kNN, and the LSA followed by the kNN. Then, the combining method shows almost same accuracies as SVM.