Uniform Partitioning of Data Grid for Association Detection

Uniform Partitioning of Data Grid for Association Detection
复制标题

DOI:
10.1109/tpami.2020.3029487
复制
发表时间:
2022-02-01
影响因子:
23.6
通讯作者:
Baraniuk, Richard G.
Baraniuk, Richard G.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Mousavi, Ali;Baraniuk, Richard G.

文献摘要

被引文献

相似文献

从大型数据集中推断出适当的信息已经变得很重要。特别是,识别这些数据集中变量之间的关系具有深远的影响。在本文中,我们介绍了统一的信息系数(UIC),该信息系数(UIC)衡量了两个多维变量之间的依赖性量,并能够检测线性和非线性关联。我们提出的UIC受到最大信息系数(MIC)[1]的启发。但是,麦克风最初设计用于测量两个一维变量之间的依赖性。与取决于两个变量之间关联类型的MIC计算不同,我们表明,UIC计算在计算上较不昂贵,并且对两个变量之间的关联类型更健壮。 UIC通过基于数据网格的均匀分区来替换MIC计算中的动态编程步骤来实现这一目标。该计算效率是以不最大化麦克风算法完成的信息系数的成本。我们为UIC的性能和各种实验提供了理论保证,以证明其在检测关联方面的质量。
Inferring appropriate information from large datasets has become important. In particular, identifying relationships among variables in these datasets has far-reaching impacts. In this article, we introduce the uniform information coefficient (UIC), which measures the amount of dependence between two multidimensional variables and is able to detect both linear and non-linear associations. Our proposed UIC is inspired by the maximal information coefficient (MIC) [1].; however, the MIC was originally designed to measure dependence between two one-dimensional variables. Unlike the MIC calculation that depends on the type of association between two variables, we show that the UIC calculation is less computationally expensive and more robust to the type of association between two variables. The UIC achieves this by replacing the dynamic programming step in the MIC calculation with a simpler technique based on the uniform partitioning of the data grid. This computational efficiency comes at the cost of not maximizing the information coefficient as done by the MIC algorithm. We present theoretical guarantees for the performance of the UIC and a variety of experiments to demonstrate its quality in detecting associations.