Testing Conditional Independence of Discrete Distributions

Testing Conditional Independence of Discrete Distributions
复制标题

DOI:
10.1145/3188745.3188756
复制
发表时间:
2017-11
期刊:
2018 Information Theory and Applications Workshop (ITA)
影响因子:
--
通讯作者:
C. Canonne;Ilias Diakonikolas;D. Kane;Alistair Stewart
C. Canonne;Ilias Diakonikolas;D. Kane;Alistair Stewart
中科院分区:
其他
文献类型:
--
作者:
C. Canonne;Ilias Diakonikolas;D. Kane;Alistair Stewart

文献摘要

相似文献

我们研究了离散分布的有条件独立性的问题。 n] $,我们想以至少2/3的概率区分,在$ x $和$ y $与$ z $的情况下,$ x $和$ y $与(x,\ y,\ z)为ɛ的情况下是独立的-far,在ℓ1距离,从每个具有此分布的分布财产。有条件的独立性是在各种科学领域中具有一系列应用的概率和统计学的核心概念。世纪。复杂性是已知的,即使对于重要的特殊情况,X和Y的域是二进制的。时间[ell_ {2}] \ times [n] $。具体而言,对于典型设置时,当ℓ1,ℓ2= o(1)时,我们表明测试有条件独立性的样本复杂性(上限和匹配下限)为\ begin {equination*} \ theta(\ theta(\ max(n^{{n^{) 1/2}/\ varepsilon^{2},\ min(n^{7/8}/varepsilon,n^{6/7}/\ varepsilon^{8/7}))为了获得我们的测试仪,我们采用了各种工具,包括(1)适当的加权改编[DK16]和(2)针对以下独立兴趣的统计问题的最佳(无偏)估计器的设计和分析:给定学位-D多项式Q:RN→R和样本访问分布$ P $的访问[ n],Q(p_ {1},…,p_ {n}),直至小添加性错误。 ]这项工作的贡献是,我们开发了一个通用理论,为所有此类估计量提供了紧密的差异,并使用相互信息方法建立,依赖于可能在其他设置中有用的硬实例的新结构。
We study the problem of testing conditional independence for discrete distributions. Specifically, given samples from a discrete random variable (X, Y, Z) on domain $[\ell_{1}]\times[\ell_{2}]\times[n]$, we want to distinguish, with probability at least 2/3, between the case that $X$ and $Y$ are conditionally independent given $Z$ from the case that (X,\ Y,\ Z) is ɛ-far, in ℓ1-distance, from every distribution that has this property. Conditional independence is a concept of central importance in probability and statistics with a range of applications in various scientific domains. As such, the statistical task of testing conditional independence has been extensively studied in various forms within the statistics and econometrics communities for nearly a century. Perhaps surprisingly, this problem has not been previously considered in the framework of distribution property testing and in particular no tester with sublinear sample complexity is known, even for the important special case that the domains of X and Y are binary. The main algorithmic result of this work is the first conditional independence tester with sublinear sample complexity for discrete distributions over $[\ell_{1}]\times[ell_{2}]\times[n]$. To complement our upper bounds, we prove information-theoretic lower bounds establishing that the sample complexity of our algorithm is optimal, up to constant factors, for a number of settings. Specifically, for the prototypical setting when ℓ1,ℓ2=O(1), we show that the sample complexity of testing conditional independence (upper bound and matching lower bound) is \begin{equation*} \Theta(\max(n^{1/2}/\varepsilon^{2},\min(n^{7/8}/varepsilon,n^{6/7}/\varepsilon^{8/7}))) \end{equation*} To obtain our tester, we employ a variety of tools, including (1) a suitable weighted adaptation of the flattening technique [DK16], and (2) the design and analysis of an optimal (unbiased) estimator for the following statistical problem of independent interest: Given a degree -d polynomial Q:Rn→R and sample access to a distribution $p$ over [n], estimate Q(p_{1}, …, p_{n}) up to small additive error. Obtaining tight variance analyses for specific estimators of this form has been a major technical hurdle in distribution testing (see, e.g., [CDVV14]). As an important contribution of this work, we develop a general theory providing tight variance bounds for all such estimators. Our lower bounds, established using the mutual information method, rely on novel constructions of hard instances that may be useful in other settings.