High-dimensional causal discovery under non-Gaussianity

High-dimensional causal discovery under non-Gaussianity
复制标题

DOI:
10.1093/biomet/asz055
复制
发表时间:
2018-03
期刊:
影响因子:
2.7
通讯作者:
Y Samuel Wang;Mathias Drton
Y Samuel Wang;Mathias Drton
中科院分区:
数学2区
文献类型:
--
作者:
Y Samuel Wang;Mathias Drton

文献摘要

被引文献

相似文献

我们考虑基于线性结构方程递归系统的图形模型。这意味着存在变量的排序$\sigma$,使得每个观察变量$Y_v$是变量特定误差项和其他观察变量$Y_u$与$\sigma(U)<\sigma(V)$的线性函数。因果关系,即线性函数依赖于哪些其他变量,可以使用有向图来描述。以前已经证明,当变量特定的误差项是非高斯时,可以从观测数据一致地估计确切的因果图,而不是马尔可夫等价类。我们提出了一种算法,在高维环境下也可以得到对图的一致估计,在高维环境中,变量的数量可能以比观测数量更快的速度增长,但在其中潜在的因果结构具有适当的稀疏性;具体地说,图的最大程度是受控制的。我们的理论分析是在对数凹误差分布的设置下进行的。
We consider graphical models based on a recursive system of linear structural equations. This implies that there is an ordering, $\sigma$, of the variables such that each observed variable $Y_v$ is a linear function of a variable-specific error term and the other observed variables $Y_u$ with $\sigma(u) < \sigma (v)$. The causal relationships, i.e., which other variables the linear functions depend on, can be described using a directed graph. It has previously been shown that when the variable-specific error terms are non-Gaussian, the exact causal graph, as opposed to a Markov equivalence class, can be consistently estimated from observational data. We propose an algorithm that yields consistent estimates of the graph also in high-dimensional settings in which the number of variables may grow at a faster rate than the number of observations, but in which the underlying causal structure features suitable sparsity; specifically, the maximum in-degree of the graph is controlled. Our theoretical analysis is couched in the setting of log-concave error distributions.