Grouped Heterogeneous Mixture Modeling for Clustered Data

Grouped Heterogeneous Mixture Modeling for Clustered Data
复制标题

DOI:
10.1080/01621459.2020.1777136
复制
发表时间:
2018-04
影响因子:
3.7
通讯作者:
S. Sugasawa
S. Sugasawa
中科院分区:
数学1区
文献类型:
--
作者:
S. Sugasawa

文献摘要

被引文献

相似文献

摘要在科学研究的各个领域中,数据的加密是普遍存在的。在这篇文章中,我们提出了一个灵活的和可解释的建模方法,称为分组异构混合建模,聚类数据,模型集群明智的条件分布的混合物的潜在的条件分布共同的所有集群。在该模型中,我们假设集群被划分为有限数量的组和混合比例是相同的在同一组内。我们提供了一个简单的广义EM算法计算的最大似然估计,和信息准则来选择组和潜在分布的数量。我们还提出了结构化的分组策略,通过在似然函数中引入分组参数的惩罚。在聚类数和聚类大小都趋于无穷大的情况下,给出了极大似然估计和信息准则的渐近性质。我们证明了所提出的方法,通过模拟研究和应用在东京的犯罪风险建模。
Abstract Clustered data are ubiquitous in a variety of scientific fields. In this article, we propose a flexible and interpretable modeling approach, called grouped heterogeneous mixture modeling, for clustered data, which models cluster-wise conditional distributions by mixtures of latent conditional distributions common to all the clusters. In the model, we assume that clusters are divided into a finite number of groups and mixing proportions are the same within the same group. We provide a simple generalized EM algorithm for computing the maximum likelihood estimator, and an information criterion to select the numbers of groups and latent distributions. We also propose structured grouping strategies by introducing penalties on grouping parameters in the likelihood function. Under the settings where both the number of clusters and cluster sizes tend to infinity, we present asymptotic properties of the maximum likelihood estimator and the information criterion. We demonstrate the proposed method through simulation studies and an application to crime risk modeling in Tokyo.