Max-Sliced Mutual Information

Max-Sliced Mutual Information
复制标题

DOI:
10.48550/arxiv.2309.16200
复制
发表时间:
2023-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Dor Tsur;Ziv Goldfeld;K. Greenewald
Dor Tsur;Ziv Goldfeld;K. Greenewald
中科院分区:
其他
文献类型:
--
作者:
Dor Tsur;Ziv Goldfeld;K. Greenewald

文献摘要

相似文献

量化高维随机变量之间的依赖性是统计学习和推理的核心。两种经典方法是典型相关分析 (CCA),它识别原始变量的最大相关投影版本;以及香农互信息,它是一种通用依赖性度量,也捕获高阶依赖性。然而,CCA 只考虑了线性相关性,这对于某些应用来说可能是不够的,而互信息通常无法在高维度上计算/估计。这项工作提出了 CCA 的可扩展信息理论推广形式的中间立场,称为最大切片互信息 (mSMI)。 mSMI 等于高维变量的低维投影之间的最大互信息,在高斯情况下还原为 CCA。它兼具两全其美的优点:捕获数据中复杂的依赖关系,同时能够从样本中进行快速计算和可扩展的估计。我们证明 mSMI 保留了香农互信息的有利结构特性,例如变分形式和独立性识别。然后,我们研究 mSMI 的统计估计,提出一种有效可计算的神经估计器,并将其与正式的非渐近误差界结合起来。我们提出的实验证明了 mSMI 在多项任务中的实用性,包括独立测试、多视图表示学习、算法公平性和生成建模。我们观察到,mSMI 始终优于竞争方法,并且几乎没有计算开销。
Quantifying the dependence between high-dimensional random variables is central to statistical learning and inference. Two classical methods are canonical correlation analysis (CCA), which identifies maximally correlated projected versions of the original variables, and Shannon's mutual information, which is a universal dependence measure that also captures high-order dependencies. However, CCA only accounts for linear dependence, which may be insufficient for certain applications, while mutual information is often infeasible to compute/estimate in high dimensions. This work proposes a middle ground in the form of a scalable information-theoretic generalization of CCA, termed max-sliced mutual information (mSMI). mSMI equals the maximal mutual information between low-dimensional projections of the high-dimensional variables, which reduces back to CCA in the Gaussian case. It enjoys the best of both worlds: capturing intricate dependencies in the data while being amenable to fast computation and scalable estimation from samples. We show that mSMI retains favorable structural properties of Shannon's mutual information, like variational forms and identification of independence. We then study statistical estimation of mSMI, propose an efficiently computable neural estimator, and couple it with formal non-asymptotic error bounds. We present experiments that demonstrate the utility of mSMI for several tasks, encompassing independence testing, multi-view representation learning, algorithmic fairness, and generative modeling. We observe that mSMI consistently outperforms competing methods with little-to-no computational overhead.