Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks

Harnessing the Power of Infinitely Wide Deep Nets on Small-data Tasks
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Sanjeev Arora;S. Du;Zhiyuan Li;R. Salakhutdinov;Ruosong Wang;Dingli Yu
Sanjeev Arora;S. Du;Zhiyuan Li;R. Salakhutdinov;Ruosong Wang;Dingli Yu
中科院分区:
其他
文献类型:
--
作者:
Sanjeev Arora;S. Du;Zhiyuan Li;R. Salakhutdinov;Ruosong Wang;Dingli Yu

文献摘要

被引文献

相似文献

最近的研究表明,以下两个模型是等效的:(a)通过梯度下降在L2损失下训练的无限宽神经网络(NNS),其无穷小的学习率(b)内核回归相对于所谓的神经切线内核(NTKS)。 (Jacot等,2018)。 Arora等人出现了一种有效的计算NTK及其卷积对应物的算法。 (2019a),它允许研究CIFAR-10等数据集上无限宽的网络的性能。但是,内核方法的超季节运行时间使其最适合小型数据任务。我们报告的结果表明,神经切线内核在低数据任务上表现出色。 1。在UCI数据库的分类/回归任务的标准测试中,NTK SVM击败了先前的金标准,随机森林(RF)以及相应的有限网。 2。在10-640训练样本的CIFAR -10上,卷积NTK始终以Resnet -34击败1%-3%。 3。在VOC07测试床上,用于通过传输学习的ImageNet上的几个图像分类任务(Goyal等,2019),替换了当前使用卷积NTK SVM使用的线性SVM始终提高性能。 4。将NTK的性能与它得出的有限宽度网的性能进行比较,NTK行为的范围比理论分析所建议的要低的净宽度(Arora等,2019a)。 NTK的功效可能会追溯到输出的差异较低。
Recent research shows that the following two models are equivalent: (a) infinitely wide neural networks (NNs) trained under l2 loss by gradient descent with infinitesimally small learning rate (b) kernel regression with respect to so-called Neural Tangent Kernels (NTKs) (Jacot et al., 2018). An efficient algorithm to compute the NTK, as well as its convolutional counterparts, appears in Arora et al. (2019a), which allowed studying performance of infinitely wide nets on datasets like CIFAR-10. However, super-quadratic running time of kernel methods makes them best suited for small-data tasks. We report results suggesting neural tangent kernels perform strongly on low-data tasks. 1. On a standard testbed of classification/regression tasks from the UCI database, NTK SVM beats the previous gold standard, Random Forests (RF), and also the corresponding finite nets. 2. On CIFAR-10 with 10 - 640 training samples, Convolutional NTK consistently beats ResNet-34 by 1% - 3%. 3. On VOC07 testbed for few-shot image classification tasks on ImageNet with transfer learning (Goyal et al., 2019), replacing the linear SVM currently used with a Convolutional NTK SVM consistently improves performance. 4. Comparing the performance of NTK with the finite-width net it was derived from, NTK behavior starts at lower net widths than suggested by theoretical analysis(Arora et al., 2019a). NTK's efficacy may trace to lower variance of output.