An integration of fuzzy association rules and WordNet for document clustering

An integration of fuzzy association rules and WordNet for document clustering
复制标题

DOI:
10.1007/s10115-010-0364-2
复制
发表时间:
2009-04
影响因子:
2.7
通讯作者:
Chun-Ling Chen;F. S. Tseng;Tyne Liang
Chun-Ling Chen;F. S. Tseng;Tyne Liang
中科院分区:
计算机科学4区
文献类型:
--
作者:
Chun-Ling Chen;F. S. Tseng;Tyne Liang

文献摘要

被引文献

相似文献

随着文本文档的快速增长,文档聚类技术应运而生,以实现高效的文档检索和更好的文档浏览。近年来,人们提出了利用关联规则挖掘中的频繁项集对文档进行聚类的方法,以解决文档的高维性、可扩展性、准确性和有意义的聚类标签等问题。为了提高文档聚类结果的质量,提出了一种有效的基于模糊频繁项集的文档聚类(F2IDC)方法,该方法将模糊关联规则挖掘与WordNet中嵌入的背景知识相结合。从WordNet中生成的术语层次结构被应用于发现广义频繁项集作为对文档进行分组的候选聚类标签。我们已经在Classic4,Re0,R8和WebKB数据集上进行了实验来评估我们的方法。我们的实验结果表明,我们提出的方法确实提供了更准确的聚类结果比以前有影响力的聚类方法在最近的文献。
With the rapid growth of text documents, document clustering technique is emerging for efficient document retrieval and better document browsing. Recently, some methods had been proposed to resolve the problems of high dimensionality, scalability, accuracy, and meaningful cluster labels by using frequent itemsets derived from association rule mining for clustering documents. In order to improve the quality of document clustering results, we propose an effective Fuzzy Frequent Itemset-based Document Clustering (F2IDC) approach that combines fuzzy association rule mining with the background knowledge embedded in WordNet. A term hierarchy generated from WordNet is applied to discover generalized frequent itemsets as candidate cluster labels for grouping documents. We have conducted experiments to evaluate our approach on Classic4, Re0, R8, and WebKB datasets. Our experimental results show that our proposed approach indeed provide more accurate clustering results than prior influential clustering methods presented in recent literature.