On similarity indices and correction for chance agreement

On similarity indices and correction for chance agreement
复制标题

DOI:
10.1007/s00357-006-0017-z
复制
发表时间:
2006-09-01
影响因子:
2
通讯作者:
Mihalko, Daniel
Mihalko, Daniel
中科院分区:
计算机科学4区
文献类型:
--
作者:
Albatineh, Ahmed N.;Niewiadomska-Bugaj, Magdalena;Mihalko, Daniel

文献摘要

被引文献

相似文献

相似性指数可用于比较数据集的分区(聚类)。多年来,文献中介绍了许多此类指数。我们显示,在我们能够跟踪的28个指数中,有22个不同的指数。即使它们的值对于所比较的相同聚类不同,在校正仅归因于机会的一致性之后,它们的值变得相似,其中一些甚至变得相等。因此,用于比较不同聚类的指数的选择问题变得不那么重要。
Similarity indices can be used to compare partitions (clusterings) of a data set. Many such indices were introduced in the literature over the years. We are showing that out of 28 indices we were able to track, there are 22 different ones. Even though their values differ for the same clusterings compared, after correcting for agreement attributed to chance only, their values become similar and some of them even become equivalent. Consequently, the problem of choice of the index to be used for comparing different clusterings becomes less important.