CCMI : Classifier based Conditional Mutual Information Estimation

CCMI : Classifier based Conditional Mutual Information Estimation
复制标题

DOI:
--
复制
发表时间:
2019-06
期刊:
--
影响因子:
--
通讯作者:
Sudipto Mukherjee;Himanshu Asnani;Sreeram Kannan
Sudipto Mukherjee;Himanshu Asnani;Sreeram Kannan
中科院分区:
其他
文献类型:
--
作者:
Sudipto Mukherjee;Himanshu Asnani;Sreeram Kannan

文献摘要

被引文献

相似文献

条件互信息 (CMI) 是在给定另一个随机变量 Z 的情况下随机变量 X 和 Y 之间的条件依赖性的度量。它可用于量化许多数据驱动的推理问题中变量之间的条件依赖性,例如图模型、因果学习、特征选择和时间序列分析。虽然基于 k 最近邻 (kNN) 的估计器以及基于核的方法已广泛用于 CMI 估计,但它们严重受到维数灾难的影响。在本文中,我们利用分类器和生成模型的进步来设计 CMI 估计方法。具体来说,我们通过训练分类器来引入基于似然比的 KL 散度估计器,以区分观察到的联合分布和乘积分布。然后,我们通过从条件生成模型中汲取灵感,展示如何使用这个基本散度估计器构建多个 CMI 估计器。我们证明,我们提出的方法的估计不会随着维度的增加而降低性能,并且比广泛使用的 KSG 估计器获得显着改进。最后,作为准确 CMI 估计的应用,我们使用最好的估计器进行条件独立性测试,并在模拟和真实数据集上实现比最先进的测试器更优异的性能。
Conditional Mutual Information (CMI) is a measure of conditional dependence between random variables X and Y, given another random variable Z. It can be used to quantify conditional dependence among variables in many data-driven inference problems such as graphical models, causal learning, feature selection and time-series analysis. While k-nearest neighbor (kNN) based estimators as well as kernel-based methods have been widely used for CMI estimation, they suffer severely from the curse of dimensionality. In this paper, we leverage advances in classifiers and generative models to design methods for CMI estimation. Specifically, we introduce an estimator for KL-Divergence based on the likelihood ratio by training a classifier to distinguish the observed joint distribution from the product distribution. We then show how to construct several CMI estimators using this basic divergence estimator by drawing ideas from conditional generative models. We demonstrate that the estimates from our proposed approaches do not degrade in performance with increasing dimension and obtain significant improvement over the widely used KSG estimator. Finally, as an application of accurate CMI estimation, we use our best estimator for conditional independence testing and achieve superior performance than the state-of-the-art tester on both simulated and real data-sets.