EAGER: Quantifying the error landscape of deep neural networks
EAGER: Quantifying the error landscape of deep neural networks
批准号:
2226387
负责人:
Stefano Martiniani
金额:
$14.92万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-02-15 至 2024-08-31
中文摘要
深度学习系统在广泛的应用中取得的显著成功可以归因于它们很好地近似复杂函数的能力,它们被有效训练的能力,以及它们在预测未知输入值方面的良好表现。最后一个性质,即泛化,特别令人费解。可以观察到,由随机梯度下降优化算法训练的深度神经网络(dnn)产生的模型泛化效果很好,特别是当模型参数的数量大大超过模型所训练的样本数量时。传统的理论无法解释这些现象,需要新的研究视角和方法来阐明这些现象。为此,统计力学可能提供能够解决深度学习中长期存在的问题的方法和观点。能量格局代表了这些领域交叉点的一个共同范式:当训练深度神经网络时,我们将所谓的“误差格局”下降到与特定模型参数选择相对应的最小值。理解dnn的泛化性能就等于理解误差结构和训练算法的动态变化之间的相互作用。特别是,“平坦最小值”的概念作为对这些观察结果的可能解释而越来越受欢迎,但是缺乏一种严格的估计平坦度的方法。我们建议采用统计力学中开发的一类新方法来回答有关dnn误差景观结构的问题,并确定找到给定解的概率,其平坦度和泛化性能之间的关系。这条调查路线应该对我们对深度学习系统泛化的理解产生重大影响,并对交通、安全和医疗等高风险应用产生影响。该建议旨在为深度神经网络的误差景观特征以及景观结构和优化动力学之间的相互作用如何产生可推广的解决方案带来新的严格程度。因此,我们将能够阐明为什么dnn具有低估计误差(即高泛化性能)。这样的理解将代表着深度学习理论发展的重要一步。我们的目标是利用最先进的科学数值技术来测量高维参数空间中吸引力盆地的体积。我们将测量流域体积分布和相关平坦度作为参数数量和网络泛化性能的函数。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
The remarkable success achieved by deep learning systems in a broad number of applications can be attributed to their ability to approximate complex functions well, their aptitude to being trained efficiently, and their good performance in predicting the values of unseen inputs. This last property, known as generalization, is particularly puzzling. It is observed that deep neural networks (DNNs) trained by the optimization algorithm known as stochastic gradient descent produce models that generalize well, particularly when the number of model parameters greatly exceeds the number of samples on which the model is trained. Traditional theory fails to explain these observations and new perspectives and means of investigation are necessary to elucidate these phenomena. To this end, statistical mechanics may provide methods and perspectives capable of addressing long-standing questions in deep learning. The energy landscape represents a common paradigm at the intersection of these fields: when training a DNN we descend the so- called “error landscape” towards a minimum corresponding to a particular choice of model parameters. Understanding generalization performance in DNNs amounts to understanding the interplay between the structure of the error landscape and the dynamics of the training algorithm that descends it. In particular, the concept of “flat minima” is gaining popularity as a possible explanation for these observations, but a rigorous approach for estimating flatness is lacking. We propose to employ a new class of methods developed within statistical mechanics to answer questions concerning the structure of the error landscapes of DNNs and to identify the relationship between the probability of finding a given solution, its flatness and its generalization performance. This line of investigation should have a significant impact on our understanding of generalization in deep learning systems with implications for high-stakes applications such as transportation, security and medicine.This proposal seeks to bring a new degree of rigor in the characterization of the error landscape of DNNs and how the interplay between landscape structure and optimization dynamics yield generalizable solutions. As a result, we will be able to elucidate why DNNs are endowed with low estimation error (i.e., high generalization performance). Such an understanding will represent a significant step forward in the development of a theory of deep learning. We aim to do so by exploiting state-of-the-science numerical techniques to measure the volume of basins of attraction in high-dimensional parameter spaces. We will measure the basin volume distributions and the associated flatness as a function of the number of parameters and the generalization performance of the network.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1063/5.0137111
发表时间:
2023-01-28
期刊:
JOURNAL OF CHEMICAL PHYSICS
影响因子:
4.4
作者:
[Anzivino, Carmine, Casiulis, Mathias, Zaccone, Alessio]
通讯作者:
Zaccone, Alessio
GOALI: Frameworks: At-Scale Heterogeneous Data based Adaptive Development Platform for Machine-Learning Models for Material and Chemical Discovery
-
批准号:2311632
-
项目类别:Standard Grant
-
资助金额:$450.0万
-
财政年份:2023
-
负责人:Stefano Martiniani
-
依托单位:
EAGER: Quantifying the error landscape of deep neural networks
-
批准号:2132995
-
项目类别:Standard Grant
-
资助金额:$14.92万
-
财政年份:2021
-
负责人:Stefano Martiniani
-
依托单位:
海外基金