CAREER: Learning from demonstrations and beyond -- consolidating imitation and reinforcement learning
CAREER: Learning from demonstrations and beyond -- consolidating imitation and reinforcement learning
批准号:
2238979
负责人:
Guni Sharon
金额:
$58.45万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-06-01 至 2028-05-31
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Recent advancements in deep reinforcement learning (RL) hold unprecedented potential for automating and optimizing control of real-world tasks such as autonomous driving, traffic management, medical procedures, robotic manufacturing, and energy management. Unfortunately, it is common for RL algorithms to exhibit unstable and/or inefficient learning, which limits their applicability. Seeking to address this critical concern, this CAREER project leverages imitation learning (IL), or behavior copying, which is better understood and typically more stable. The project targets the unification of IL and RL into a holistic paradigm that can safely and effectively learn from, and outperform, existing solutions. This project will address outstanding knowledge gaps in both types of learning through a novel curriculum decomposition of the tasks, where simplified demonstrations are used to bootstrap the learner’s behavior. The project will also foster education and outreach activities. Specifically, it will enhance undergraduate STEM training by providing students with exposure to scientific research and knowledge discovery processes relating to safety-critical AI applications through an original multidisciplinary undergraduate engineering program. Moreover, it will facilitate a unique K12 outreach activity within a large minority (Hispanic/latino) community (Bryan, TX). The project will support and advance an existing research collaboration with an industrial partner in the context of defense technology. This collaboration, in turn, is expected to advance the US national defense.This project will form the basis for a new research thrust in ML---one that combines IL and RL toward a holistic, robust, and safe learning framework. It will define and prove a no-regret bound on the training process within the Markov-Decision Process formalization. The approach is to reduce an IL problem to an RL one that includes a domain-independent curriculum-learning trajectory. The resulting algorithms and solutions are expected to achieve state-of-the-art performance in complex control domains as well as to deepen theoretical understanding of the potential and limitations of the resulting solutions. Specifically, the research seeks to prove conditions guaranteeing policy convergence and monotonic improvement during training. Moreover, the project will develop domain-specific adaptation to and analysis of real-world applications (autonomous driving and robotics testbeds) while providing stable and efficient RL from demonstrations.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Comparison between popular Genetic Algorithm (GA)-based tool and Covariance Matrix Adaptation - Evolutionary Strategy (CMA-ES) for optimizing indoor daylight
用于优化室内日光的流行的基于遗传算法 (GA) 的工具与协方差矩阵适应 - 进化策略 (CMA-ES) 的比较
DOI:
10.26868/25222708.2023.1218
发表时间:
2023
期刊:
Proceedings of Building Simulation 2023: 18th Conference of IBPSA
影响因子:
--
作者:
[Anis, Manal, Pendurkar, Sumedh, Yi, Yun Kyu, Sharon, Guni]
通讯作者:
Sharon, Guni
The (Un)Scalability of Informed Heuristic Function Estimation in NP-Hard Search Problems
NP 难搜索问题中知情启发式函数估计的(非)可扩展性
DOI:
--
发表时间:
2023
期刊:
Transactions on Machine Learning Research
影响因子:
--
作者:
[Sumedh Pendurkar, Taoan Huang, Brendan Juba, Jiapeng Zhang, Sven Koenig, Guni Sharon]
通讯作者:
Guni Sharon
DOI:
10.5555/3545946.3598887
发表时间:
2023
期刊:
影响因子:
--
作者:
[Sumedh Pendurkar;Chris Chow;Luo Jie;Guni Sharon]
通讯作者:
Sumedh Pendurkar;Chris Chow;Luo Jie;Guni Sharon
Task Phasing: Automated Curriculum Learning from Demonstrations
任务阶段化:从演示中自动进行课程学习
DOI:
10.1609/icaps.v33i1.27235
发表时间:
2023
期刊:
Proceedings of the International Conference on Automated Planning and Scheduling
影响因子:
--
作者:
[Bajaj, Vaibhav, Sharon, Guni, Stone, Peter]
通讯作者:
Stone, Peter
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于有向超图的大型个性化e-learning学习过程模型的自动生成与优化
-
批准号:61572533
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:孙雪冬
-
依托单位:
E-Learning中学习者情感补偿方法的研究
-
批准号:61402392
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2014
-
负责人:秦继伟
-
依托单位: