Document clustering based on constructing density tree?

Document clustering based on constructing density tree?
复制标题

DOI:
10.1007/s12209-008-0005-y
复制
发表时间:
2008-02
影响因子:
7.1
通讯作者:
Weidi Dai;Wenjun Wang;Yuexian Hou;Ying Wang;Lu Zhang
Weidi Dai;Wenjun Wang;Yuexian Hou;Ying Wang;Lu Zhang
中科院分区:
--
文献类型:
--
作者:
Weidi Dai;Wenjun Wang;Yuexian Hou;Ying Wang;Lu Zhang

文献摘要

相似文献

为了提高文档聚类的准确率,研究了基于DEnsityTree的文档聚类算法(CABDET)。CABDET方法根据局部密度动态调整邻域半径,为每个潜在簇构建基于密度的树结构。它避免了基于密度的带噪声应用程序的空间聚集(DBSCAN)的S全局密度参数,并将输入参数减少到1。在真实文档上的实验结果表明,CABDET比DBSCAN方法获得了更高的聚类精度。当邻域半径为0.80时,CABDET算法得到的最大F-度量值为0.347,高于邻域半径为0.65、最小对象数为6的DBSCAN算法的0.332。
This paper focuses on document clustering by clustering algorithm based on a DEnsityTree (CABDET) to improve the accuracy of clustering. The CABDET method constructs a density-based treestructure for every potential cluster by dynamically adjusting the radius of neighborhood according to local density. It avoids density-based spatial clustering of applications with noise (DBSCAN)’s global density parameters and reduces input parameters to one. The results of experiment on real document show that CABDET achieves better accuracy of clustering than DBSCAN method. The CABDET algorithm obtains the maxF-measure value 0.347 with the root node’s radius of neighborhood 0.80, which is higher than 0.332 of DBSCAN with the radius of neighborhood 0.65 and the minimum number of objects 6.