CAREER: Understanding the Inductive Biases in Modern Machine Learning
CAREER: Understanding the Inductive Biases in Modern Machine Learning
批准号:
1943251
负责人:
Raman Arora
金额:
$50.0万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-02-15 至 2025-01-31
中文摘要
现代机器学习(特别是深度学习)的最新进展正在引领人工智能时代,这有可能彻底改变我们日常生活的方方面面。然而,就像蒸汽机的早期一样,对深度学习的令人满意的理解到目前为止还难以捉摸。我们目前缺乏一个正式的深度学习理论,这个理论可以解释为什么我们可以用看似不够的训练数据训练过于复杂的模型,并且仍然找到推广到以前看不见的数据的解决方案,或者为什么为一个任务训练的模型在另一个相关任务上也表现良好,或者为什么训练过的模型如此容易受到轻微的、几乎难以察觉的数据损坏的影响。该项目旨在通过开发一种与实践紧密结合并受到实践推动的解释性和规定性深度学习理论来解决这一需求。该方法不是将学习简单地视为一个黑盒优化问题,而是通过揭示算法启发式来研究内部工作原理,算法启发式在赋予训练模型出色的泛化特性方面可能发挥同样重要的作用。鉴于深度学习的广泛适用性以及拟议研究中理论分析和实证研究的互补性,该项目特别适合将研究整合到教育和推广中。拟议的教育活动包括课程开发、暑期实习、黑客马拉松以及通过巴尔的摩当地项目开展的教师拓展活动。该项目研究了显式算法正则化在早期停止、批处理规范化和辍学形式中的作用,以及优化算法和网络架构的选择,以提供有助于泛化的适当归纳偏差。该项目的第二个首要目标是更广泛地理解深度学习中的泛化现象。它试图理解为什么记忆训练数据的系统仍然可以很好地泛化,神经网络架构如何实现迁移学习,以及如何设计健壮的算法,以保证深度学习解决方案在数据对抗性破坏的情况下泛化。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Recent advances in modern machine learning (deep learning in particular) are ushering in the era of artificial intelligence, which has the potential to revolutionize every aspect of our daily lives. However, much like the early days of the steam engine, a satisfactory understanding of deep learning has so far been elusive. We currently lack a formal theory of deep learning, one that could explain why we can train overly complex models with seemingly not enough training data and still find solutions that generalize to previously unseen data, or why models trained for one task also perform well on another related task, or why trained models are so vulnerable to slight, nearly imperceptible, corruptions of data. This project aims to address this need by developing an explanatory and prescriptive theory of deep learning that is tightly integrated with and motivated by the practice. Rather than view learning as simply a black-box optimization problem, the approach investigates the inner workings by shedding light on algorithmic heuristics that potentially play an equally important role in endowing the trained models with excellent generalization properties. Given the broad applicability of deep learning and the complementary nature of theoretical analyses and empirical studies in the proposed research, the project is particularly suited for integrating research into education and outreach. The proposed educational activities include curriculum development, summer internships, hackathons, and instructor's outreach through local Baltimore programs. The project investigates the role of explicit algorithmic regularization in the form of early stopping, batch normalization, and dropout, as well as the choice of optimization algorithms and network architecture in providing an adequate inductive bias that helps with generalization. A second overarching goal of the project is to understand, more broadly, the generalization phenomenon in deep learning. It seeks to understand why systems that memorize the training data can still generalize well, how the neural network architecture enables transfer learning, and how to design robust algorithms that will guarantee that deep learning solutions generalize despite adversarial corruption to data.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(23)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2024
期刊:
Advances in neural information processing systems
影响因子:
--
作者:
[Poorya Mianjy, Raman Arora]
通讯作者:
Poorya Mianjy, Raman Arora
Differentially Private Generalized Linear Models Revisited
重新审视差分私有广义线性模型
DOI:
--
发表时间:
2022
期刊:
Advances in neural information processing systems
影响因子:
--
作者:
[Raman Arora, Raef Bassily]
通讯作者:
Raman Arora, Raef Bassily
DOI:
--
发表时间:
2020-07
期刊:
ArXiv
影响因子:
--
作者:
[D. Rothchild;Ashwinee Panda;Enayat Ullah;Nikita Ivkin;I. Stoica;Vladimir Braverman;Joseph Gonzalez-Joseph-Gonzale]
通讯作者:
D. Rothchild;Ashwinee Panda;Enayat Ullah;Nikita Ivkin;I. Stoica;Vladimir Braverman;Joseph Gonzalez-Joseph-Gonzale
DOI:
--
发表时间:
2023
期刊:
Transactions on machine learning research
影响因子:
--
作者:
[Enayat Ullah, Raman Arora]
通讯作者:
Enayat Ullah, Raman Arora
On Instance-Dependent Bounds for Offline Reinforcement Learning with Linear Function Approximation
基于线性函数逼近的离线强化学习的实例相关界限
DOI:
--
发表时间:
2023
期刊:
Proceedings of the AAAI Conference on Artificial Intelligence
影响因子:
--
作者:
[Nguyễn-Tang, Thanh, Yin, Ming, Gupta, Sunil, Venkatesh, Svetha, Arora, Raman]
通讯作者:
Arora, Raman
共 21 条
BIGDATA: F: Privacy in Unsupervised Learning
-
批准号:1838139
-
项目类别:Standard Grant
-
资助金额:$91.14万
-
财政年份:2018
-
负责人:Raman Arora
-
依托单位:
BIGDATA: Collaborative Research: F: Stochastic Approximation for Subspace and Multiview Representation Learning
-
批准号:1546482
-
项目类别:Standard Grant
-
资助金额:$70.47万
-
财政年份:2015
-
负责人:Raman Arora
-
依托单位:
国内基金
海外基金
Navigating Sustainability: Understanding Environm ent,Social and Governanc e Challenges and Solution s for Chinese Enterprises
in Pakistan's CPEC Framew
ork
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:Noshaba Aziz
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
Understanding complicated gravitational physics by simple two-shell systems
-
批准号:12005059
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:国分隆文
-
依托单位: