Identification of Linear Latent Variable Model with Arbitrary Distribution

Identification of Linear Latent Variable Model with Arbitrary Distribution
复制标题

DOI:
10.1609/aaai.v36i6.20585
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Zhen Chen;Feng Xie;Jie Qiao;Z. Hao;Kun Zhang;Ruichu Cai
Zhen Chen;Feng Xie;Jie Qiao;Z. Hao;Kun Zhang;Ruichu Cai
中科院分区:
其他
文献类型:
--
作者:
Zhen Chen;Feng Xie;Jie Qiao;Z. Hao;Kun Zhang;Ruichu Cai

文献摘要

相似文献

跨多个学科的一个重要问题是推断和理解有意义的潜在变量。一种常用的策略是在关于从潜伏期到被测量的连通性的适当假设下,根据潜在变量对被测量变量进行建模(称为测量模型)。此外,发现潜在变量之间的因果关系(称为结构模型)可能更有趣。最近,人们提出了一些方法来估计结构模型,这些方法假定测量变量和潜在变量中的噪声项是非高斯的。然而,当一些噪声项变成高斯时,它们就不适用了。为了弥补这一差距,我们研究了具有任意噪声分布的结构模型的辨识问题。我们给出了结构模型可辨识的充要条件:对于每一对相邻的潜在变量Lx,Ly,(1)Lx和Ly中至少有一个具有非高斯噪声,或者(2)它们中至少有一个具有非高斯祖先,并且不因Lx和Ly的共同原因而与该祖先的非高斯分量d-分离。这种可辨识性结果将非高斯性要求放宽到只有一个(希望是很小的)变量子集,并相应地优雅地扩展了结构模型的应用范围。基于上述可识别性结果,我们进一步提出了一种实用的结构模型学习算法。通过实证研究,验证了可辨识性结果的正确性和该方法的有效性。
An important problem across multiple disciplines is to infer and understand meaningful latent variables. One strategy commonly used is to model the measured variables in terms of the latent variables under suitable assumptions on the connectivity from the latents to the measured (known as measurement model). Furthermore, it might be even more interesting to discover the causal relations among the latent variables (known as structural model). Recently, some methods have been proposed to estimate the structural model by assuming that the noise terms in the measured and latent variables are non-Gaussian. However, they are not suitable when some of the noise terms become Gaussian. To bridge this gap, we investigate the problem of identification of the structural model with arbitrary noise distributions. We provide necessary and sufficient condition under which the structural model is identifiable: it is identifiable iff for each pair of adjacent latent variables Lx, Ly, (1) at least one of Lx and Ly has non-Gaussian noise, or (2) at least one of them has a non-Gaussian ancestor and is not d-separated from the non-Gaussian component of this ancestor by the common causes of Lx and Ly. This identifiability result relaxes the non-Gaussianity requirements to only a (hopefully small) subset of variables, and accordingly elegantly extends the application scope of the structural model. Based on the above identifiability result, we further propose a practical algorithm to learn the structural model. We verify the correctness of the identifiability result and the effectiveness of the proposed method through empirical studies.