A comparative study of TF*IDF, LSI and multi-words for text classification

A comparative study of TF*IDF, LSI and multi-words for text classification
复制标题

DOI:
10.1016/j.eswa.2010.08.066
复制
发表时间:
2011-03-01
影响因子:
8.5
通讯作者:
Tang, Xijin
Tang, Xijin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhang, Wen;Yoshida, Taketoshi;Tang, Xijin

文献摘要

被引文献

相似文献

文本表示是文本挖掘的主要研究内容之一,是基于文本的智能信息处理的基础和不可缺少的部分。一般来说,文本表示包括两个任务:索引和加权。本文对TF*IDF、LSI和多词文本表示进行了比较研究。我们使用一个中文和一个英文文档集分别对这三种方法在信息检索和文本分类中的应用进行了评价。实验结果表明,在文本分类中,LSI在两个文档集上都具有比其他方法更好的性能。此外,LSI在检索英文文档方面的表现最好。这一结果表明,LSI具有良好的语义和统计质量,并与声称LSI不能产生区分力的索引不同。(C)2010爱思唯尔有限公司版权所有。
One of the main themes in text mining is text representation, which is fundamental and indispensable for text-based intellegent information processing. Generally, text representation inludes two tasks: indexing and weighting. This paper has comparatively studied TF*IDF, LSI and multi-word for text representation. We used a Chinese and an English document collection to respectively evaluate the three methods in information retreival and text categorization. Experimental results have demonstrated that in text categorization, LSI has better performance than other methods in both document collections. Also, LSI has produced the best performance in retrieving English documents. This outcome has shown that LSI has both favorable semantic and statistical quality and is different with the claim that LSI can not produce discriminative power for indexing. (C) 2010 Elsevier Ltd. All rights reserved.