Cluster detection and clustering with random start forward searches

Cluster detection and clustering with random start forward searches
复制标题

DOI:
10.1080/02664763.2017.1310806
复制
发表时间:
2018-01-01
影响因子:
1.5
通讯作者:
Cerioli, Andrea
Cerioli, Andrea
中科院分区:
数学4区
文献类型:
--
作者:
Atkinson, Anthony C.;Riani, Marco;Cerioli, Andrea

文献摘要

被引文献

相似文献

前向搜索是一种稳健的数据分析方法,其中在模型拟合中使用不断增加的数据的无离群值子集;然后根据与模型的接近程度对数据进行排序。这里,具有许多随机起点的前向搜索用于对多元数据进行聚类。这些随机开始导致对暂定簇的诊断识别。将前向搜索应用于所提出的各个簇,通过将非簇成员识别为外围成员来建立簇成员资格。该方法不需要有关聚类数量的先验信息,并且不寻求对所有观察结果进行分类。对瑞士纸币的 200 个六维观察结果的分析说明了这些特性。说明了链接图和刷涂在阐明数据结构中的重要性。我们还提供了一种自动确定聚类中心的方法,并将我们的方法的行为与基于模型的聚类进行比较。在具有八个聚类的模拟示例中,我们的方法提供了比基于模型的聚类更稳定、更准确的解决方案。我们考虑这两个过程的计算要求。
The forward search is a method of robust data analysis in which outlier free subsets of the data of increasing size are used in model fitting; the data are then ordered by closeness to the model. Here the forward search, with many random starts, is used to cluster multivariate data. These random starts lead to the diagnostic identification of tentative clusters. Application of the forward search to the proposed individual clusters leads to the establishment of cluster membership through the identification of non-cluster members as outlying. The method requires no prior information on the number of clusters and does not seek to classify all observations. These properties are illustrated by the analysis of 200 six-dimensional observations on Swiss banknotes. The importance of linked plots and brushing in elucidating data structures is illustrated. We also provide an automatic method for determining cluster centres and compare the behaviour of our method with model-based clustering. In a simulated example with eight clusters our method provides more stable and accurate solutions than model-based clustering. We consider the computational requirements of both procedures.