AN EXAMINATION OF PROCEDURES FOR DETERMINING THE NUMBER OF CLUSTERS IN A DATA SET

AN EXAMINATION OF PROCEDURES FOR DETERMINING THE NUMBER OF CLUSTERS IN A DATA SET
复制标题

DOI:
10.1007/bf02294245
复制
发表时间:
1985-01-01
期刊:
影响因子:
3
通讯作者:
COOPER, MC
COOPER, MC
中科院分区:
心理学4区
文献类型:
--
作者:
MILLIGAN, GW;COOPER, MC

文献摘要

被引文献

相似文献

30个程序的Monte Carlo评估,确定集群的数量进行人工数据集,其中包含2,3,4,或5个不同的非重叠集群。为了提供多种聚类解决方案,数据集进行了分析,通过四个层次聚类方法。外部标准的措施表明优秀的恢复真正的集群结构的方法在正确的层次结构。因此,数据中存在的聚类相当强。停止规则的模拟结果显示,它们在确定数据中正确聚类数的能力上有很大的差异。有几个程序运行得相当好,而其他程序运行得相当差。因此,后一组规则似乎没有什么有效性,特别是对于包含不同聚类的数据集。应用研究人员被敦促选择一个或多个更好的标准。然而,用户应注意,某些标准的执行可能取决于数据。
A Monte Carlo evaluation of 30 procedures for determining the number of clusters was conducted on artificial data sets which contained either 2, 3, 4, or 5 distinct nonoverlapping clusters. To provide a variety of clustering solutions, the data sets were analyzed by four hierarchical clustering methods. External criterion measures indicated excellent recovery of the true cluster structure by the methods at the correct hierarchy level. Thus, the clustering present in the data was quite strong. The simulation results for the stopping rules revealed a wide range in their ability to determine the correct number of clusters in the data. Several procedures worked fairly well, whereas others performed rather poorly. Thus, the latter group of rules would appear to have little validity, particularly for data sets containing distinct clusters. Applied researchers are urged to select one or more of the better criteria. However, users are cautioned that the performance of some of the criteria may be data dependent.