Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate Reduction

Learning Diverse and Discriminative Representations via the Principle of Maximal Coding Rate Reduction
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Yaodong Yu;Kwan Ho Ryan Chan;Chong You;Chaobing Song;Yi Ma
Yaodong Yu;Kwan Ho Ryan Chan;Chong You;Chaobing Song;Yi Ma
中科院分区:
其他
文献类型:
--
作者:
Yaodong Yu;Kwan Ho Ryan Chan;Chong You;Chaobing Song;Yi Ma

文献摘要

相似文献

为了从最能区分类别的高维数据中学习内在的低维结构,我们提出了最大编码率降低($\text{MCR}^2$)的原则,这是一种信息论度量,可以最大化整个数据集与每个单独类别之和之间的编码率差异。阐明了它与交叉熵、信息瓶颈、信息增益、收缩学习和对比学习等现有框架的关系,为学习多样性和区分性特征提供了理论保证。编码速率可以从退化子空间类分布的有限样本中准确计算,并且可以以统一的方式在监督、自监督和无监督设置中学习内在表示。从经验上讲,单独使用这一原则学习的表示比使用交叉熵的表示对分类中的标记损坏更鲁棒,并且可以导致从自学的不变特征聚类混合数据的最新结果。
To learn intrinsic low-dimensional structures from high-dimensional data that most discriminate between classes, we propose the principle of Maximal Coding Rate Reduction ($\text{MCR}^2$), an information-theoretic measure that maximizes the coding rate difference between the whole dataset and the sum of each individual class. We clarify its relationships with most existing frameworks such as cross-entropy, information bottleneck, information gain, contractive and contrastive learning, and provide theoretical guarantees for learning diverse and discriminative features. The coding rate can be accurately computed from finite samples of degenerate subspace-like distributions and can learn intrinsic representations in supervised, self-supervised, and unsupervised settings in a unified manner. Empirically, the representations learned using this principle alone are significantly more robust to label corruptions in classification than those using cross-entropy, and can lead to state-of-the-art results in clustering mixed data from self-learned invariant features.