Research on K-means Text Clustering Algorithm Based on Semantic

Research on K-means Text Clustering Algorithm Based on Semantic
复制标题

基于语义的K-means文本聚类算法研究

DOI:
10.1109/ccie.2010.39
复制
发表时间:
2010
期刊:
2010 International Conference on Computing, Control and Industrial Engineering
影响因子:
--
通讯作者:
Shuicai Shi
Shuicai Shi
中科院分区:
--
文献类型:
--
作者:
Yufang Liu;Shibin Xiao;Xueqiang Lv;Shuicai Shi

文献摘要

被引文献

相似文献

通过对文本聚类的K-Means算法和基于语义的向量空间模型的研究,提出了一种基于语义的K-Means文本聚类模型,以解决文本数据集的高维稀疏特性。该模型减少了文本数据的语义损失,提高了文本聚类的质量。实验证明,在F1指标值的最终评估中,基于语义的文本聚类比基于非语义的文本聚类提高了6%以上。
Through research on K-means algorithm of text clustering and semantic-based vector space model, a semantic-based K-means text clustering model is proposed to solve the problem on high-dimensional and sparse characteristics of text data set. The model reduces the semantic loss of the text data and improves the quality of text clustering. Experiments prove that semantic-based text clustering increases by more 6 percent than non-semantic-based one in the final evaluation of the F1 index value.