课题基金 / 基金详情

CAREER: Learning from demonstrations and beyond -- consolidating imitation and reinforcement learning

CAREER: Learning from demonstrations and beyond -- consolidating imitation and reinforcement learning
职业:从演示中学习以及超越——巩固模仿和强化学习
批准号:
2238979
负责人:
Guni Sharon
金额:
$58.45万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2028-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
深度强化学习(RL)的最新进展在自动驾驶、交通管理、医疗程序、机器人制造和能源管理等现实世界任务的自动化和优化控制方面具有前所未有的潜力。不幸的是,RL算法经常表现出不稳定和/或低效的学习,这限制了它们的适用性。为了解决这一关键问题,这个职业项目利用了模仿学习(IL),或行为复制,它更容易被理解,通常也更稳定。该项目的目标是将IL和RL统一到一个可以安全有效地学习并超越现有解决方案的整体范例中。这个项目将通过一种新的课程分解任务来解决这两种类型学习中的突出知识差距,其中简化的演示被用来引导学习者的行为。该项目还将促进教育和外联活动。具体地说,它将通过原创的多学科本科工程项目,为学生提供与安全关键的人工智能应用相关的科学研究和知识发现过程,从而加强本科STEM培训。此外,它还将促进在一个庞大的少数族裔(西班牙裔/拉丁裔)社区(德克萨斯州布赖恩)开展一项独特的K12外联活动。该项目将支持和推进与一个工业伙伴在国防技术方面的现有研究合作。这一合作预计将促进美国的国防。该项目将为ML的新研究推力奠定基础-将IL和RL结合起来,建立一个全面、稳健和安全的学习框架。它将在马尔可夫决策过程形式化中定义并证明训练过程的无悔界。方法是将IL问题简化为包含独立于领域的课程学习轨迹的RL问题。由此产生的算法和解决方案有望在复杂的控制领域实现最先进的性能,并加深对最终解决方案的潜力和局限性的理论理解。具体地说,这项研究试图证明保证政策收敛和培训期间单调改进的条件。此外,该项目将开发特定领域的适应和分析真实世界的应用程序(自动驾驶和机器人测试床),同时提供稳定和高效的演示RL。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Recent advancements in deep reinforcement learning (RL) hold unprecedented potential for automating and optimizing control of real-world tasks such as autonomous driving, traffic management, medical procedures, robotic manufacturing, and energy management. Unfortunately, it is common for RL algorithms to exhibit unstable and/or inefficient learning, which limits their applicability. Seeking to address this critical concern, this CAREER project leverages imitation learning (IL), or behavior copying, which is better understood and typically more stable. The project targets the unification of IL and RL into a holistic paradigm that can safely and effectively learn from, and outperform, existing solutions. This project will address outstanding knowledge gaps in both types of learning through a novel curriculum decomposition of the tasks, where simplified demonstrations are used to bootstrap the learner’s behavior. The project will also foster education and outreach activities. Specifically, it will enhance undergraduate STEM training by providing students with exposure to scientific research and knowledge discovery processes relating to safety-critical AI applications through an original multidisciplinary undergraduate engineering program. Moreover, it will facilitate a unique K12 outreach activity within a large minority (Hispanic/latino) community (Bryan, TX). The project will support and advance an existing research collaboration with an industrial partner in the context of defense technology. This collaboration, in turn, is expected to advance the US national defense.This project will form the basis for a new research thrust in ML---one that combines IL and RL toward a holistic, robust, and safe learning framework. It will define and prove a no-regret bound on the training process within the Markov-Decision Process formalization. The approach is to reduce an IL problem to an RL one that includes a domain-independent curriculum-learning trajectory. The resulting algorithms and solutions are expected to achieve state-of-the-art performance in complex control domains as well as to deepen theoretical understanding of the potential and limitations of the resulting solutions. Specifically, the research seeks to prove conditions guaranteeing policy convergence and monotonic improvement during training. Moreover, the project will develop domain-specific adaptation to and analysis of real-world applications (autonomous driving and robotics testbeds) while providing stable and efficient RL from demonstrations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
Comparison between popular Genetic Algorithm (GA)-based tool and Covariance Matrix Adaptation - Evolutionary Strategy (CMA-ES) for optimizing indoor daylight
用于优化室内日光的流行的基于遗传算法 (GA) 的工具与协方差矩阵适应 - 进化策略 (CMA-ES) 的比较
DOI: 10.26868/25222708.2023.1218
发表时间: 2023
期刊: Proceedings of Building Simulation 2023: 18th Conference of IBPSA
影响因子: --
作者: [Anis, Manal, Pendurkar, Sumedh, Yi, Yun Kyu, Sharon, Guni]
通讯作者: Sharon, Guni
The (Un)Scalability of Informed Heuristic Function Estimation in NP-Hard Search Problems
NP 难搜索问题中知情启发式函数估计的(非)可扩展性
DOI: --
发表时间: 2023
期刊: Transactions on Machine Learning Research
影响因子: --
作者: [Sumedh Pendurkar, Taoan Huang, Brendan Juba, Jiapeng Zhang, Sven Koenig, Guni Sharon]
通讯作者: Guni Sharon
DOI: 10.5555/3545946.3598887
发表时间: 2023
期刊:
影响因子: --
作者: [Sumedh Pendurkar;Chris Chow;Luo Jie;Guni Sharon]
通讯作者: Sumedh Pendurkar;Chris Chow;Luo Jie;Guni Sharon
Task Phasing: Automated Curriculum Learning from Demonstrations
任务阶段化:从演示中自动进行课程学习
DOI: 10.1609/icaps.v33i1.27235
发表时间: 2023
期刊: Proceedings of the International Conference on Automated Planning and Scheduling
影响因子: --
作者: [Bajaj, Vaibhav, Sharon, Guni, Stone, Peter]
通讯作者: Stone, Peter
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: