COMPARING PARTITIONS

COMPARING PARTITIONS
复制标题

DOI:
10.1007/bf01908075
复制
发表时间:
1985-01-01
影响因子:
2
通讯作者:
ARABIE, P
ARABIE, P
中科院分区:
计算机科学4区
文献类型:
--
作者:
HUBERT, L;ARABIE, P

文献摘要

被引文献

相似文献

比较有限对象集的两个不同分区的问题在集群文献中不断出现。我们首先回顾了Rand(1971)通常认为的一个著名的划分对应度量,讨论了为机会修正这个指标的问题,并注意到最近由Morey和Gonsti(1984)提出并被其他人(例如Miligan和Cooper 1985)采用的归一化策略基于一个不正确的假设。然后,通过使用一个简单的叉积度量来评估两个邻近矩阵的一致性,从而间接地解决了比较划分的一般问题。它们是使用各种评分规则从相应的分区生成的。可派生的特殊情况包括传统上熟悉的统计和/或为不同地加权某些对象对而定制的统计。最后,我们提出了一种基于对象三元组的比较的度量方法,除了具有概率解释的优点外,还可以进行偶然性校正(即在合理的零假设下假设一个常量值),并在±1之间有界。
The problem of comparing two different partitions of a finite set of objects reappears continually in the clustering literature. We begin by reviewing a well-known measure of partition correspondence often attributed to Rand (1971), discuss the issue of correcting this index for chance, and note that a recent normalization strategy developed by Morey and Agresti (1984) and adopted by others (e.g., Miligan and Cooper 1985) is based on an incorrect assumption. Then, the general problem of comparing partitions is approached indirectly by assessing the congruence of two proximity matrices using a simple cross-product measure. They are generated from corresponding partitions using various scoring rules. Special cases derivable include traditionally familiar statistics and/or ones tailored to weight certain object pairs differentially. Finally, we propose a measure based on the comparison of object triples having the advantage of a probabilistic interpretation in addition to being corrected for chance (i.e., assuming a constant value under a reasonable null hypothesis) and bounded between ±1.