An Evaluation on Feature Selection for Text Clustering

An Evaluation on Feature Selection for Text Clustering
复制标题

DOI:
--
复制
发表时间:
2003-08
期刊:
--
影响因子:
--
通讯作者:
Tao Liu;Shengping Liu;Zheng Chen;Wei-Ying Ma
Tao Liu;Shengping Liu;Zheng Chen;Wei-Ying Ma
中科院分区:
其他
文献类型:
--
作者:
Tao Liu;Shengping Liu;Zheng Chen;Wei-Ying Ma

文献摘要

被引文献

相似文献

特征选择方法已成功应用于文本分类,但由于类别标签信息的不可用,很少应用于文本聚类。在本文中,我们首先给出了特征选择方法可以提高文本聚类算法的效率和性能的经验证据。然后,我们提出了一种称为“术语贡献(TC)”的新特征选择方法,并对文本聚类的多种特征选择方法进行了比较研究,包括文档频率(DF)、术语强度(TS)、基于熵(En)、信息增益(IG)和χ2统计(CHI)。最后,我们提出了一种“迭代特征选择(IF)”方法,通过利用有效的监督特征选择方法迭代选择特征并执行聚类来解决标签不可用的问题。论文中提供了 Web 目录数据的详细实验结果。
Feature selection methods have been successfully applied to text categorization but seldom applied to text clustering due to the unavailability of class label information. In this paper, we first give empirical evidence that feature selection methods can improve the efficiency and performance of text clustering algorithm. Then we propose a new feature selection method called "Term Contribution (TC)" and perform a comparative study on a variety of feature selection methods for text clustering, including Document Frequency (DF), Term Strength (TS), Entropy-based (En), Information Gain (IG) and χ2 statistic (CHI). Finally, we propose an "Iterative Feature Selection (IF)" method that addresses the unavailability of label problem by utilizing effective supervised feature selection method to iteratively select features and perform clustering. Detailed experimental results on Web Directory data are provided in the paper.