Learning with invariances in random features and kernel models

Learning with invariances in random features and kernel models
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Song Mei;Theodor Misiakiewicz;A. Montanari
Song Mei;Theodor Misiakiewicz;A. Montanari
中科院分区:
其他
文献类型:
--
作者:
Song Mei;Theodor Misiakiewicz;A. Montanari

文献摘要

被引文献

相似文献

许多机器学习任务需要高度的不变性:如果我们对数据进行特定的一组变换,数据分布不会改变。例如,图像的标签在图像的平移下是不变的。某些神经网络架构-例如卷积网络-被认为是将其成功归功于它们利用这种不变性的事实。为了量化不变结构所获得的增益,我们引入了两类模型:不变随机特征和不变核方法。作为特殊情况,后者包括具有全局平均池化的卷积网络的神经切线内核。我们考虑球面和超立方体上的一致协变量分布和一般不变目标函数。我们刻画了高维区域中不变方法的测试误差,其中样本大小和隐藏单元的数量在维度上按多项式缩放,对于一类我们称之为“退化群”的群,α ≤ 1。我们发现,利用不变性的架构节省了d因子(d代表的尺寸)的样本大小和隐藏单元的数量,以实现相同的测试误差为非结构化架构。最后,我们表明,输出对称化的非结构化核估计并没有得到显着的统计改善;另一方面,数据增强与非结构化核估计是等价的不变核估计,并享有相同的统计效率的改善。
A number of machine learning tasks entail a high degree of invariance: the data distribution does not change if we act on the data with a certain group of transformations. For instance, labels of images are invariant under translations of the images. Certain neural network architectures —for instance, convolutional networks—are believed to owe their success to the fact that they exploit such invariance properties. With the objective of quantifying the gain achieved by invariant architectures, we introduce two classes of models: invariant random features and invariant kernel methods. The latter includes, as a special case, the neural tangent kernel for convolutional networks with global average pooling. We consider uniform covariates distributions on the sphere and hypercube and a general invariant target function. We characterize the test error of invariant methods in a high-dimensional regime in which the sample size and number of hidden units scale as polynomials in the dimension, for a class of groups that we call ‘degeneracy α’, with α ≤ 1. We show that exploiting invariance in the architecture saves a d factor (d stands for the dimension) in sample size and number of hidden units to achieve the same test error as for unstructured architectures. Finally, we show that output symmetrization of an unstructured kernel estimator does not give a significant statistical improvement; on the other hand, data augmentation with an unstructured kernel estimator is equivalent to an invariant kernel estimator and enjoys the same improvement in statistical efficiency.