Mixture model selection via hierarchical BIC

Mixture model selection via hierarchical BIC
复制标题

通过分层 BIC 进行混合模型选择

DOI:
10.1016/j.csda.2015.01.019
复制
发表时间:
2015
影响因子:
1.8
通讯作者:
Lei Shi
Lei Shi
中科院分区:
数学3区
文献类型:
--
作者:
Jianhua Zhao;Libin Jin;Lei Shi

文献摘要

相似文献

贝叶斯信息准则 (BIC) 是有限混合模型中最流行的模型选择准则之一。然而,它使用整个样本量来惩罚每个组件的复杂性,并完全忽略了数据固有的聚类结构,导致过度惩罚。为了克服这个问题,提出了一种称为分层 BIC(HBIC)的新标准,该标准仅使用其局部样本大小来惩罚组件复杂性,并与集群数据结构很好地匹配。理论上,当样本量较大时,HBIC 是变分贝叶斯 (VB) 下界的近似值,而广泛使用的 BIC 是不太精确的近似值。我们进行了实证研究来验证这一理论结果,并在模拟和真实数据集上进行了一系列实验来比较 HBIC 和 BIC。结果显示,HBIC 的表现大幅优于 BIC,而 BIC 则遭受低估。
The Bayesian information criterion (BIC) is one of the most popular criteria for model selection in finite mixture models. However, it implausibly penalizes the complexity of each component using the whole sample size and completely ignores the clustered structure inherent in the data, resulting in over-penalization. To overcome this problem, a novel criterion called hierarchical BIC (HBIC) is proposed which penalizes the component complexity only using its local sample size and matches the clustered data structure well. Theoretically, HBIC is an approximation of the variational Bayesian (VB) lower bound when sample size is large and the widely used BIC is a less accurate approximation. An empirical study is conducted to verify this theoretical result and a series of experiments is performed on simulated and real data sets to compare HBIC and BIC. The results show that HBIC outperforms BIC substantially and BIC suffers from underestimation.