A Distribution Free Conditional Independence Test with Applications to Causal Discovery

A Distribution Free Conditional Independence Test with Applications to Causal Discovery
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Zhanrui Cai;Runze Li;Yaowu Zhang
Zhanrui Cai;Runze Li;Yaowu Zhang
中科院分区:
其他
文献类型:
--
作者:
Zhanrui Cai;Runze Li;Yaowu Zhang

文献摘要

相似文献

本文主要研究条件独立性的检验问题。我们首先建立了条件独立性和相互独立性之间的等价关系。基于等价性,我们提出了一个指标来衡量的条件依赖,通过量化的相互依赖的转换变量。建议的索引有几个吸引人的属性。(a)它是分布自由的,因为所提出的指数的极限零分布不依赖于数据的总体分布。因此,临界值可以通过模拟制成表格。(b)所提出的索引范围从0到1,等于零当且仅当条件独立性成立。因此,它在备择假设下具有非平凡的功效。(c)它是强大的离群值和重尾数据,因为它是不变的条件严格单调变换。(d)它具有较低的计算成本,因为它采用了一个简单的封闭形式的表达式,可以在二次时间内实现。(e)它是不敏感的调整参数的计算所提出的指数。(f)新的指数适用于多元随机向量以及离散数据。所有这些性质使我们能够使用新的指数作为统计推断工具的各种数据。通过大量的仿真和真实的因果发现的应用,说明了该方法的有效性。
This paper is concerned with test of the conditional independence. We first establish an equivalence between the conditional independence and the mutual independence. Based on the equivalence, we propose an index to measure the conditional dependence by quantifying the mutual dependence among the transformed variables. The proposed index has several appealing properties. (a) It is distribution free since the limiting null distribution of the proposed index does not depend on the population distributions of the data. Hence the critical values can be tabulated by simulations. (b) The proposed index ranges from zero to one, and equals zero if and only if the conditional independence holds. Thus, it has nontrivial power under the alternative hypothesis. (c) It is robust to outliers and heavy-tailed data since it is invariant to conditional strictly monotone transformations. (d) It has low computational cost since it incorporates a simple closed-form expression and can be implemented in quadratic time. (e) It is insensitive to tuning parameters involved in the calculation of the proposed index. (f) The new index is applicable for multivariate random vectors as well as for discrete data. All these properties enable us to use the new index as statistical inference tools for various data. The effectiveness of the method is illustrated through extensive simulations and a real application on causal discovery.