Validation of overlapping clustering: A random clustering perspective

Validation of overlapping clustering: A random clustering perspective
复制标题

DOI:
10.1016/j.ins.2010.07.028
复制
发表时间:
2010-11
期刊:
Inf. Sci.
影响因子:
--
通讯作者:
Junjie Wu;Hua Yuan;Hui Xiong;Guoqing Chen
Junjie Wu;Hua Yuan;Hui Xiong;Guoqing Chen
中科院分区:
其他
文献类型:
--
作者:
Junjie Wu;Hua Yuan;Hui Xiong;Guoqing Chen

文献摘要

被引文献

相似文献

f测度作为一种应用广泛的聚类验证测度,在信息检索领域受到越来越多的关注。在本文中,我们揭示了当f测度用于验证具有不同簇数(增量效应)或相关文档的不同先验概率(先验概率效应)的数据时,它可能导致对重叠簇结果的偏见观点。本文从随机聚类的角度出发,在f测度的基础上提出了一种新的“隐含强度”测度。此外,我们仔细研究了IMI的性质。最后,在真实数据集上的实验结果表明,IMI显著缓解了f测度固有的偏增量效应和先验概率效应。
As a widely used clustering validation measure, the F-measure has received increased attention in the field of information retrieval. In this paper, we reveal that the F-measure can lead to biased views as to results of overlapped clusters when it is used for validating the data with different cluster numbers (incremental effect) or different prior probabilities of relevant documents (prior-probability effect). We propose a new “IMplication Intensity” (IMI) measure which is based on the F-measure and is developed from a random clustering perspective. In addition, we carefully investigate the properties of IMI. Finally, experimental results on real-world data sets show that IMI significantly alleviates biased incremental and prior-probability effects which are inherent to the F-measure.