Margins of discrete Bayesian networks

Margins of discrete Bayesian networks
复制标题

DOI:
10.1214/17-aos1631
复制
发表时间:
2015-01
期刊:
The Annals of Statistics
影响因子:
--
通讯作者:
R. Evans
R. Evans
中科院分区:
其他
文献类型:
--
作者:
R. Evans

文献摘要

相似文献

具有潜在变量的贝叶斯网络模型在统计学和机器学习中有着广泛的应用。在本文中,我们提供了一个完整的代数特征的贝叶斯网络模型的潜变量时,观察变量是离散的,没有假设的状态空间的潜变量。我们证明了它在代数上等价于所谓的嵌套马尔可夫模型,这意味着两者在联合概率上的不等式约束是相同的。特别是这两个模型具有相同的尺寸。因此,嵌套马尔可夫模型是潜在变量模型的最佳可能描述,避免了考虑一般极其复杂的不等式。这样做的结果是Tian和Pearl(UAI 2002,pp 519 -527)的约束查找算法对于查找等式约束是完备的。潜变量模型存在参数不可识别和非正则渐近的困难;相反,嵌套马尔可夫模型是完全可识别的,表示已知维度的弯曲指数族,并且可以很容易地使用显式参数化进行拟合。
Bayesian network models with latent variables are widely used in statistics and machine learning. In this paper we provide a complete algebraic characterization of Bayesian network models with latent variables when the observed variables are discrete and no assumption is made about the state-space of the latent variables. We show that it is algebraically equivalent to the so-called nested Markov model, meaning that the two are the same up to inequality constraints on the joint probabilities. In particular these two models have the same dimension. The nested Markov model is therefore the best possible description of the latent variable model that avoids consideration of inequalities, which are extremely complicated in general. A consequence of this is that the constraint finding algorithm of Tian and Pearl (UAI 2002, pp519-527) is complete for finding equality constraints. Latent variable models suffer from difficulties of unidentifiable parameters and non-regular asymptotics; in contrast the nested Markov model is fully identifiable, represents a curved exponential family of known dimension, and can easily be fitted using an explicit parameterization.