Measuring Dependence Powerfully and Equitably

Measuring Dependence Powerfully and Equitably
复制标题

DOI:
--
复制
发表时间:
2015-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yakir A. Reshef;David N. Reshef;H. Finucane;Pardis C Sabeti;M. Mitzenmacher
Yakir A. Reshef;David N. Reshef;H. Finucane;Pardis C Sabeti;M. Mitzenmacher
中科院分区:
其他
文献类型:
--
作者:
Yakir A. Reshef;David N. Reshef;H. Finucane;Pardis C Sabeti;M. Mitzenmacher

文献摘要

被引文献

相似文献

给定一个高维数据集,我们通常希望找到其中最强的关系。一个常见的策略是评估每个变量对的依赖性度量,并保留得分最高的变量对进行后续操作。如果使用的统计数据是公平的[Reshef et al. 2015 a],即,如果对于噪声的某种度量,它将类似的分数分配给同等噪声的关系而不管关系类型(例如,线性的、指数的、周期的)。在本文中,我们引入并描述了一种称为MIC* 的人口依赖性度量。我们展示了MIC* 可以被视为三种方式:作为MIC的总体值,来自[Reshef et al. 2011]的高度公平的统计量,作为互信息的规范“平滑”,以及作为根据联合分布边缘的最佳一维分区定义的无限序列的上确界。基于这一理论,我们介绍了一种有效的方法来计算MIC* 从一对随机变量的密度,我们定义了一个新的一致的估计MIC *,是有效的计算。相比之下,没有已知的多项式时间算法来计算原始的公平统计MIC。我们通过模拟表明,MICe比MIC具有更好的偏差-方差特性。然后,我们介绍并证明了第二个统计,TICe,这是一个微不足道的副产品的计算MICe和其目标是强大的独立性测试,而不是公平的一致性。我们在模拟中表明,MICe和TICe有很好的公平性和权力反对独立性。这里的分析补充了对几种主要依赖性指标的更深入的实证评估[Reshef et al. 2015 b],显示了MICe和TICe的最新性能。
Given a high-dimensional data set we often wish to find the strongest relationships within it. A common strategy is to evaluate a measure of dependence on every variable pair and retain the highest-scoring pairs for follow-up. This strategy works well if the statistic used is equitable [Reshef et al. 2015a], i.e., if, for some measure of noise, it assigns similar scores to equally noisy relationships regardless of relationship type (e.g., linear, exponential, periodic). In this paper, we introduce and characterize a population measure of dependence called MIC*. We show three ways that MIC* can be viewed: as the population value of MIC, a highly equitable statistic from [Reshef et al. 2011], as a canonical "smoothing" of mutual information, and as the supremum of an infinite sequence defined in terms of optimal one-dimensional partitions of the marginals of the joint distribution. Based on this theory, we introduce an efficient approach for computing MIC* from the density of a pair of random variables, and we define a new consistent estimator MICe for MIC* that is efficiently computable. In contrast, there is no known polynomial-time algorithm for computing the original equitable statistic MIC. We show through simulations that MICe has better bias-variance properties than MIC. We then introduce and prove the consistency of a second statistic, TICe, that is a trivial side-product of the computation of MICe and whose goal is powerful independence testing rather than equitability. We show in simulations that MICe and TICe have good equitability and power against independence respectively. The analyses here complement a more in-depth empirical evaluation of several leading measures of dependence [Reshef et al. 2015b] that shows state-of-the-art performance for MICe and TICe.