A SIMPLE MEASURE OF CONDITIONAL DEPENDENCE

A SIMPLE MEASURE OF CONDITIONAL DEPENDENCE
复制标题

DOI:
10.1214/21-aos2073
复制
发表时间:
2021-12-01
影响因子:
4.5
通讯作者:
Chatterjee, Sourav
Chatterjee, Sourav
中科院分区:
数学1区
文献类型:
--
作者:
Azadkia, Mona;Chatterjee, Sourav

文献摘要

被引文献

相似文献

在给定一组其他变量X-1,…,X-p的情况下,基于I.I.D.,我们提出了两个随机变量Y和Z之间的条件依赖系数。样本。该系数具有一长串理想的性质,其中最重要的是在绝对没有分布假设的情况下,它收敛于[0,1]中的一个极限,其中极限为0当且仅当Y和Z在给定X-1,…,X-p时是条件独立的,当且仅当Y等于给定X-1,…,X-p的Z的一个可测函数。此外,它有一个自然的解释,作为常见的部分R-2统计量的非线性推广,用于通过回归来衡量条件相关性。利用这个统计量,我们设计了一种新的变量选择算法,称为按条件独立性的特征排序(FOCI),该算法不需要模型,没有调整参数,并且在稀疏性假设下可证明是一致的。对合成数据集和真实数据集进行了大量的应用。
We propose a coefficient of conditional dependence between two random variables Y and Z given a set of other variables X-1, ..., X-p, based on an i.i.d. sample. The coefficient has a long list of desirable properties, the most important of which is that under absolutely no distributional assumptions, it converges to a limit in [0, 1], where the limit is 0 if and only if Y and Z are conditionally independent given X-1, ..., X-p, and is 1 if and only if Y is equal to a measurable function of Z given X-1, ..., X-p. Moreover, it has a natural interpretation as a nonlinear generalization of the familiar partial R-2 statistic for measuring conditional dependence by regression. Using this statistic, we devise a new variable selection algorithm, called Feature Ordering by Conditional Independence (FOCI), which is model-free, has no tuning parameters, and is provably consistent under sparsity assumptions. A number of applications to synthetic and real data sets are worked out.