课题基金 / 基金详情

CAREER: Overparameterization in modern machine learning: A panacea or a pitfall?

CAREER: Overparameterization in modern machine learning: A panacea or a pitfall?
职业:现代机器学习中的过度参数化:万能药还是陷阱?
批准号:
2239151
负责人:
Vidya Muthukumar
金额:
$58.46万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2028-06-30

项目摘要

项目成果

Vidya Muthukumar的其他基金

相似基金

相关文献

中文摘要
翻译
深度神经网络压倒性地主导着经验机器学习领域。然而,它们最先进的性能仍然不为人所知、脆弱且需要大量资源才能获得。特别是,它们良好的泛化特性,或者它们对以前看不见的数据做出准确预测的能力,在很大程度上是无法解释的。特别不寻常的是,与经典的机器学习模型相比,最先进的神经网络经常过度参数化;也就是说,比他们的训练数据集“大”得多。最近的研究揭示了对这种过度参数化可能带来的好处的更好理解,但仅在初级模型族中。在表现出复杂和独特行为的深度神经网络中,过度参数化的后果存在许多未知因素。在缺乏第一原理理论的情况下,深度神经网络中突出的故障模式仍然没有得到缓解,或者需要花费不必要的成本来解决,并且架构选择是以一种浪费的试错方式进行的,包括重复的训练和测试循环。这限制了深度学习技术发挥其全部潜力,特别是在高风险和资源有限的应用中。该项目将通过跨越信号处理、信息理论和在线决策的多种数学技术,弥合最近的过参数化线性模型理论与现实世界神经网络之间的差距。具体而言,该项目将:1)研究过参数化对深度神经网络测试回归和分类性能的影响;2)表征过参数化模型(线性和非线性)对对抗性扰动和数据分布显著变化的鲁棒性;3)为现代机器学习中数据驱动的模型选择设计健壮的原则。最终,该项目旨在建立基本的数学原理,不仅可以解释现代机器学习的成功推广,还可以解释其故障模式,从而为开发高效和有原则的解决方案铺平道路。该项目还将在高中和本科阶段创建和传播基础信号处理、机器学习和数据科学的教育资源,这些资源是上述研究的基础和补充。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Deep neural networks overwhelmingly dominate the empirical machine learning landscape. Their state-of-the-art performance, however, remains poorly understood, brittle, and resource-intensive to obtain. In particular, their good generalization properties, or their ability to make accurate predictions on previously unseen data, are largely unexplained. Especially unusual is that, in contrast to classical machine learning models, state-of-the-art neural networks are frequently heavily overparameterized; that is, much “larger” than their training data set. Recent research has revealed a better understanding of the possible benefits of such overparameterization, but only in elementary model families. The ramifications of overparameterization in deep neural networks, which exhibit complex and distinct behaviors, present many unknowns. In the absence of a first-principles theory, outstanding failure modes in deep neural networks remain unmitigated or unnecessarily costly to solve, and architecture selection is conducted in a wasteful trial-and-error manner that involves repeated train-and-test cycles. This limits deep learning technology from reaching its full potential, particularly in high-stakes and resource-limited applications.This project will bridge the gap between the recent theory of overparameterized linear models and real-world neural networks through a diversity of mathematical techniques spanning signal processing, information theory, and online decision-making. In particular, the project will: 1) examine the implications of overparameterization on the test regression and classification performance of deep neural networks; 2) characterize the robustness of overparameterized models (both linear and nonlinear) to adversarial perturbations and significant shifts in the data distribution; and 3) design robust principles for data-driven model selection in modern machine learning. Ultimately, this project aims to establish foundational mathematical principles to explain not only the successful generalization of modern machine learning, but also its failure modes---in turn paving the way for developing efficient and principled solutions. This project will also create and disseminate educational resources at the high school and undergraduate levels on elementary signal processing, machine learning, and data science that underlie and complement the described research.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CIF: RI: Medium: Design principles and theory for data augmentation
  • 批准号:
    2212182
  • 项目类别:
    Standard Grant
  • 资助金额:
    $120.0万
  • 财政年份:
    2022
  • 负责人:
    Vidya Muthukumar
  • 依托单位:
海外基金