Residual Similarity Based Conditional Independence Test and Its Application in Causal Discovery

Residual Similarity Based Conditional Independence Test and Its Application in Causal Discovery
复制标题

DOI:
10.1609/aaai.v36i5.20539
复制
发表时间:
2022-06
期刊:
--
影响因子:
--
通讯作者:
Hao Zhang;Shuigeng Zhou;Kun Zhang;J. Guan
Hao Zhang;Shuigeng Zhou;Kun Zhang;J. Guan
中科院分区:
其他
文献类型:
--
作者:
Hao Zhang;Shuigeng Zhou;Kun Zhang;J. Guan

文献摘要

相似文献

近年来,许多基于回归的条件独立性(CI)检验方法被提出来解决因果发现问题。这些方法通过首先从两个目标变量中去除控制集的信息,然后测试相应的残差Res1和Res2之间的独立性,提供了测试CI的替代方案。当残差线性不相关时,它们之间的独立性检验是非平凡的。由于能够在高维空间中计算内积,通常使用基于核的方法来实现这一目标,但仍然需要花费相当多的时间。本文研究了线性非高斯结构方程模型下两个线性组合之间的独立性。我们发现,这两个残差之间的依赖性,可以捕获的相似性之间的差异(Res1,Res2)和(Res1,Res3)(Res3是由随机置换)在高维空间。在此基础上,我们设计了一种新的CI检验方法SCIT,通过排列检验来控制I类错误率。所提出的方法是简单的,但更有效和有效的比现有的。当应用于因果发现时,所提出的方法在速度和II型错误率方面都优于同行,特别是在小样本的情况下,这是通过我们在各种数据集上的大量实验来验证的。
Recently, many regression based conditional independence (CI) test methods have been proposed to solve the problem of causal discovery. These methods provide alternatives to test CI by first removing the information of the controlling set from the two target variables, and then testing the independence between the corresponding residuals Res1 and Res2. When the residuals are linearly uncorrelated, the independence test between them is nontrivial. With the ability to calculate inner product in high-dimensional space, kernel-based methods are usually used to achieve this goal, but still consume considerable time. In this paper, we investigate the independence between two linear combinations under linear non-Gaussian structural equation model. We show that the dependence between the two residuals can be captured by the difference between the similarity of (Res1, Res2) and that of (Res1, Res3) (Res3 is generated by random permutation) in high-dimensional space. With this result, we design a new method called SCIT for CI test, where permutation test is performed to control Type I error rate. The proposed method is simpler yet more efficient and effective than the existing ones. When applied to causal discovery, the proposed method outperforms the counterparts in terms of both speed and Type II error rate, especially in the case of small sample size, which is validated by our extensive experiments on various datasets.