CAREER: Overparameterization in modern machine learning: A panacea or a pitfall?
CAREER: Overparameterization in modern machine learning: A panacea or a pitfall?
批准号:
2239151
负责人:
Vidya Muthukumar
金额:
$58.46万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2028-06-30
中文摘要
深度神经网络在经验机器学习领域占据压倒性的主导地位。然而,他们最先进的性能仍然鲜为人知,很脆弱,而且需要大量资源。特别是,它们良好的泛化特性,或者它们对以前未见过的数据做出准确预测的能力,在很大程度上是无法解释的。特别不同寻常的是,与经典的机器学习模型相比,最先进的神经网络经常严重过度参数化;也就是说,比它们的训练数据集“大得多”。最近的研究表明,人们对这种过度参数化可能带来的好处有了更好的理解,但仅限于初级模型家庭。深度神经网络中的过度参数化衍生物表现出复杂而独特的行为,呈现出许多未知数。在缺乏第一原理理论的情况下,深度神经网络中的突出故障模式仍然无法缓解或解决起来成本高得不必要,并且体系结构选择是以一种浪费的试错方式进行的,这涉及重复的训练和测试周期。这限制了深度学习技术充分发挥其潜力,特别是在高风险和资源有限的应用中。该项目将通过横跨信号处理、信息论和在线决策的多种数学技术,弥合最近的过度参数化线性模型理论和现实世界神经网络之间的差距。特别是,该项目将:1)研究过度参数化对深度神经网络的测试回归和分类性能的影响;2)表征过度参数化模型(线性和非线性)对对抗性扰动和数据分布的显著变化的稳健性;以及3)为现代机器学习中的数据驱动模型选择设计稳健的原则。最终,这个项目的目标是建立基本的数学原理,不仅解释现代机器学习的成功推广,而且解释其失败模式-反过来,为开发高效和有原则的解决方案铺平道路。该项目还将在高中和本科生层面创建和传播基础信号处理、机器学习和数据科学方面的教育资源,这些资源是上述研究的基础和补充。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Deep neural networks overwhelmingly dominate the empirical machine learning landscape. Their state-of-the-art performance, however, remains poorly understood, brittle, and resource-intensive to obtain. In particular, their good generalization properties, or their ability to make accurate predictions on previously unseen data, are largely unexplained. Especially unusual is that, in contrast to classical machine learning models, state-of-the-art neural networks are frequently heavily overparameterized; that is, much “larger” than their training data set. Recent research has revealed a better understanding of the possible benefits of such overparameterization, but only in elementary model families. The ramifications of overparameterization in deep neural networks, which exhibit complex and distinct behaviors, present many unknowns. In the absence of a first-principles theory, outstanding failure modes in deep neural networks remain unmitigated or unnecessarily costly to solve, and architecture selection is conducted in a wasteful trial-and-error manner that involves repeated train-and-test cycles. This limits deep learning technology from reaching its full potential, particularly in high-stakes and resource-limited applications.This project will bridge the gap between the recent theory of overparameterized linear models and real-world neural networks through a diversity of mathematical techniques spanning signal processing, information theory, and online decision-making. In particular, the project will: 1) examine the implications of overparameterization on the test regression and classification performance of deep neural networks; 2) characterize the robustness of overparameterized models (both linear and nonlinear) to adversarial perturbations and significant shifts in the data distribution; and 3) design robust principles for data-driven model selection in modern machine learning. Ultimately, this project aims to establish foundational mathematical principles to explain not only the successful generalization of modern machine learning, but also its failure modes---in turn paving the way for developing efficient and principled solutions. This project will also create and disseminate educational resources at the high school and undergraduate levels on elementary signal processing, machine learning, and data science that underlie and complement the described research.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CIF: RI: Medium: Design principles and theory for data augmentation
-
批准号:2212182
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2022
-
负责人:Vidya Muthukumar
-
依托单位:
海外基金