Anchored Bayesian Gaussian mixture models

Anchored Bayesian Gaussian mixture models
复制标题

DOI:
10.1214/20-ejs1756
复制
发表时间:
2020-01-01
影响因子:
1.1
通讯作者:
Peruggia, Mario
Peruggia, Mario
中科院分区:
数学3区
文献类型:
--
作者:
Kunkel, Deborah;Peruggia, Mario

文献摘要

被引文献

相似文献

有限混合物是一种灵活的建模工具,用于不规则形状的密度和来自异质种群的样本。当使用组分特征上的可交换先验对混合物进行建模时,组分标签是任意的,并且在后验分析中是不可区分的。这使得不可能将任何有意义的解释归因于组件特征的边缘后验分布。我们提出了一个模型,其中一个小数目的观察假设产生的一些标记的组件密度。所得到的模型是不可交换的,允许在没有后处理的情况下对组件特征进行推断。我们的方法在建模阶段为组件标签分配意义,并且可以作为标签上的数据依赖性信息先验来证明。我们表明,我们的方法产生可解释的结果,通常(但不总是)类似于重新标记算法产生的结果,额外的好处是边际推论直接来自一个指定的概率模型,而不是事后操纵。我们提供的渐近结果,导致实际的指导方针模型选择的动机是最大限度地提高先验信息的类标签,并证明我们的方法对真实的和模拟数据。
Finite mixtures are a flexible modeling tool for irregularly shaped densities and samples from heterogeneous populations. When modeling with mixtures using an exchangeable prior on the component features, the component labels are arbitrary and are indistinguishable in posterior analysis. This makes it impossible to attribute any meaningful interpretation to the marginal posterior distributions of the component features. We propose a model in which a small number of observations are assumed to arise from some of the labeled component densities. The resulting model is not exchangeable, allowing inference on the component features without post-processing. Our method assigns meaning to the component labels at the modeling stage and can be justified as a data-dependent informative prior on the labelings. We show that our method produces interpretable results, often (but not always) similar to those resulting from relabeling algorithms, with the added benefit that the marginal inferences originate directly from a well specified probability model rather than a post hoc manipulation. We provide asymptotic results leading to practical guidelines for model selection that are motivated by maximizing prior information about the class labels and demonstrate our method on real and simulated data.