Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks.

Spectral bias and task-model alignment explain generalization in kernel regression and infinitely wide neural networks.
复制标题

DOI:
10.1038/s41467-021-23103-1
复制
发表时间:
2021-05-18
影响因子:
16.6
通讯作者:
Pehlevan C
Pehlevan C
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Canatar A;Bordelon B;Pehlevan C

文献摘要

参考文献

被引文献

相似文献

对于许多机器学习模型来说,对泛化的理论理解仍然是一个悬而未决的问题,包括深度网络,其中过度参数化会带来更好的性能,这与经典统计学的传统智慧相矛盾。在这里,我们研究核回归的泛化误差,它除了是一种流行的机器学习方法之外,还描述了某些无限超参数化的神经网络。我们使用统计力学的技术来导出适用于任何内核和数据分布的泛化误差的分析表达式。我们将我们的理论应用于真实和合成数据集,以及许多内核,包括那些在无限宽度限制下训练深度网络所产生的内核。我们阐明了核回归的归纳偏差,以用简单函数解释数据,表征核是否与学习任务兼容,并表明当有噪声或核无法表达时,更多数据可能会损害泛化,导致可能有许多峰值的非单调学习曲线。卡纳塔等人。提出适用于实际数据的核回归泛化预测理论。该理论解释了在宽神经网络中观察到的各种泛化现象,这些现象承认内核限制并且尽管参数化过度但仍具有良好的泛化能力。
A theoretical understanding of generalization remains an open problem for many machine learning models, including deep networks where overparameterization leads to better performance, contradicting the conventional wisdom from classical statistics. Here, we investigate generalization error for kernel regression, which, besides being a popular machine learning method, also describes certain infinitely overparameterized neural networks. We use techniques from statistical mechanics to derive an analytical expression for generalization error applicable to any kernel and data distribution. We present applications of our theory to real and synthetic datasets, and for many kernels including those that arise from training deep networks in the infinite-width limit. We elucidate an inductive bias of kernel regression to explain data with simple functions, characterize whether a kernel is compatible with a learning task, and show that more data may impair generalization when noisy or not expressible by the kernel, leading to non-monotonic learning curves with possibly many peaks. Canatar et al. propose a predictive theory of generalization in kernel regression applicable to real data. This theory explains various generalization phenomena observed in wide neural networks, which admit a kernel limit and generalize well despite being overparameterized.
DOI: 10.1007/s102080010030
发表时间: 2002-11-01
影响因子: 3
作者:
Cucker, F;Smale, S
通讯作者: Smale, S
DOI: 10.1103/physrevx.6.031034
发表时间: 2016-08-29
期刊: PHYSICAL REVIEW X
影响因子: 12.5
作者:
Advani, Madhu;Ganguli, Surya
通讯作者: Ganguli, Surya
DOI: 10.1088/1742-5468/2005/05/p05012
发表时间: 2005-05-01
影响因子: 2.4
作者:
Castellani, T;Cavagna, A
通讯作者: Cavagna, A
DOI: 10.1073/pnas.1903070116
发表时间: 2019-08-06
影响因子: 11.1
作者:
Belkin, Mikhail;Hsu, Daniel;Mandal, Soumik
通讯作者: Mandal, Soumik
DOI: 10.1103/physrevlett.82.2975
发表时间: 1999-04-05
影响因子: 8.6
作者:
Dietrich, R;Opper, M;Sompolinsky, H
通讯作者: Sompolinsky, H