A Constructive Prediction of the Generalization Error Across Scales

A Constructive Prediction of the Generalization Error Across Scales
复制标题

DOI:
--
复制
发表时间:
2019-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Jonathan S. Rosenfeld;Amir Rosenfeld;Yonatan Belinkov;N. Shavit
Jonathan S. Rosenfeld;Amir Rosenfeld;Yonatan Belinkov;N. Shavit
中科院分区:
其他
文献类型:
--
作者:
Jonathan S. Rosenfeld;Amir Rosenfeld;Yonatan Belinkov;N. Shavit

文献摘要

被引文献

相似文献

神经网络的泛化误差对模型和数据集大小的依赖性对于实践和理解神经网络理论都至关重要。然而,这种依赖性的功能形式仍然难以捉摸。在这项工作中,我们提出了一种函数形式,可以很好地近似实践中的泛化误差。利用模型缩放(例如宽度、深度)的成功概念,我们能够同时构建这样的形式并指定可以在模型/数据尺度上实现它的确切模型。我们的构建遵循从对一系列模型/数据规模、各种模型类型和数据集、视觉和语言任务中进行的观察中获得的见解。我们表明,该形式既能很好地拟合跨尺度的观测结果,又能提供从小规模到大规模模型和数据的准确预测。
The dependency of the generalization error of neural networks on model and dataset size is of critical importance both in practice and for understanding the theory of neural networks. Nevertheless, the functional form of this dependency remains elusive. In this work, we present a functional form which approximates well the generalization error in practice. Capitalizing on the successful concept of model scaling (e.g., width, depth), we are able to simultaneously construct such a form and specify the exact models which can attain it across model/data scales. Our construction follows insights obtained from observations conducted over a range of model/data scales, in various model types and datasets, in vision and language tasks. We show that the form both fits the observations well across scales, and provides accurate predictions from small- to large-scale models and data.