Automatic selection of the number of clusters using Bayesian clustering and sparsity‐inducing priors

Automatic selection of the number of clusters using Bayesian clustering and sparsity‐inducing priors
复制标题

使用贝叶斯聚类和稀疏性诱导先验自动选择聚类数量

DOI:
10.1002/eap.2524
复制
发表时间:
2022
影响因子:
5
通讯作者:
Cullen, Joshua
Cullen, Joshua
中科院分区:
环境科学与生态学1区
文献类型:
--
作者:
Valle, Denis;Jameel, Yusuf;Betancourt, Brenda;Azeria, Ermias T.;Attias, Nina;Cullen, Joshua

文献摘要

相似文献

聚类是生态和环境科学中普遍存在的任务,并且为此目的开发了多种方法。由于这些聚类方法通常要求用户先验指定组的数量,因此标准方法是针对不同数量的组运行算法,然后使用标准(例如 AIC 或 BIC)选择最佳数量。这种方法的问题在于,多次运行这些聚类算法(即,针对不同数量的组)在计算上可能会非常昂贵,并且其中一些信息标准可能会导致对组数量的高估。为了解决这些问题,我们提倡在贝叶斯聚类框架内使用稀疏性先验。特别是,我们强调了如何使用截断棒断裂(TSB)先验(贝叶斯非参数学中常用的先验)来同时确定各种贝叶斯聚类模型的组数和估计模型参数,而无需拟合多个模型。我们使用运动生态学和群落生态学背景下的模拟数据,说明了在此之前成功恢复三种聚类模型(两种类型的混合模型,应用于 GPS 运动数据和物种出现数据,以及物种原型模型)的真实群体数量的能力。然后,我们将这些模型应用于巴西的犰狳运动数据、阿尔伯塔省(加拿大)的植物发生数据和北美的鸟类发生数据。我们相信,鉴于聚类的普遍性以及确定组数量的相关挑战,许多生态和环境科学应用将受益于具有稀疏性先验的贝叶斯聚类方法。提供了两个 R 包(EcoCluster 和 bayesmove),可以将这些模型与 TSB 先验直接拟合。
Clustering is a ubiquitous task in ecological and environmental sciences and multiple methods have been developed for this purpose. Because these clustering methods typically require users to a priori specify the number of groups, the standard approach is to run the algorithm for different numbers of groups and then choose the optimal number using a criterion (e.g., AIC or BIC). The problem with this approach is that it can be computationally expensive to run these clustering algorithms multiple times (i.e., for different numbers of groups) and some of these information criteria can lead to an overestimation of the number of groups. To address these concerns, we advocate for the use of sparsity‐inducing priors within a Bayesian clustering framework. In particular, we highlight how the truncated stick‐breaking (TSB) prior, a prior commonly adopted in Bayesian nonparametrics, can be used to simultaneously determine the number of groups and estimate model parameters for a wide range of Bayesian clustering models without requiring the fitting of multiple models. We illustrate the ability of this prior to successfully recover the true number of groups for three clustering models (two types of mixture models, applied to GPS movement data and species occurrence data, as well as the species archetype model) using simulated data in the context of movement ecology and community ecology. We then apply these models to armadillo movement data in Brazil, plant occurrence data from Alberta (Canada), and bird occurrence data from North America. We believe that many ecological and environmental sciences applications will benefit from Bayesian clustering methods with sparsity‐inducing priors given the ubiquity of clustering and the associated challenge of determining the number of groups. Two R packages, EcoCluster and bayesmove, are provided that enable the straightforward fitting of these models with the TSB prior.