S&AS: FND: Uncertainty-Aware Safe Deep Reinforcement Learning
S&AS: FND: Uncertainty-Aware Safe Deep Reinforcement Learning
批准号:
1849154
负责人:
David Held
金额:
$50.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-04-01 至 2023-03-31
中文摘要
机器人设计师不能总是预测所有现实世界的可能性;因此,机器人将需要使用能够学习适应环境意外变化的算法。然而,如果机器人要在现实世界中学习和适应,它们必须以在学习的同时保持安全的方式进行适应。该项目开发的方法可以确保机器人在学习和适应环境的意外变化时仍然安全。通过使机器人能够认识到它们的不确定性,这些方法可以帮助机器人更安全地操作。此外,这些方法将使机器人能够适应环境中不可预见的变化。这些方法将应用于自动驾驶,使机器人能够在不同摩擦水平的表面或不平坦的地形上安全操作,否则可能不安全。这些方法将使自动驾驶汽车能够在恶劣天气条件下更安全地在城市行驶,以及在越野环境中进行搜索和救援行动,巡逻车辆探测动物偷猎者以及其他应用。本研究开发了一套不确定性感知安全机器人学习方法。这些方法将使机器人能够估计其动作效果的不确定性;基于估计的不确定性,机器人将决定如何在安全谨慎操作的同时提高其性能。此外,这些方法将利用估计的不确定性来确定机器人应该如何调整其参数以有效地响应环境变化。这个项目中的方法将在复杂的策略类上运行,比如那些由深度强化学习训练的神经网络所代表的策略类。这种方法将使机器人实现安全和自适应的长期自主。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Robot designers cannot always anticipate all real-world eventualities; hence robots will need to use algorithms that can learn to adapt to unexpected changes in their environment. However, if robots are to learn and adapt in the real world, they must adapt in ways that continue to maintain safety while learning. This project develops methods that ensure that robots continue to be safe even as they learn and adapt to unexpected changes in their environment. By enabling robots to be cognizant of their uncertainty, these methods can help robots operate more safely. Further, the methods will enable robots to be adaptive to unforeseen changes in their environment. These methods will be applied to autonomous driving, to enable robots to operate safely on surfaces with different levels of friction or on uneven terrain that might otherwise be unsafe to operate on. Such methods will enable safer autonomous vehicles for city driving in poor weather conditions, as well as for operating in off-road settings for search and rescue operations, patrol vehicles to detect animal poachers, and for other applications.This research develops a set of methods for uncertainty-aware safe robot learning. These methods will enable a robot to estimate the uncertainty of the effect of its actions; based on the estimated uncertainty, the robot will determine how to improve its performance while operating safely and cautiously. Furthermore, these methods will use the estimated uncertainty to determine how the robot should adapt its parameters to efficiently respond to environmental changes. The methods in this project will operate on complex policy classes such as those represented by neural networks trained with deep reinforcement learning. Such an approach will enable robots to achieve safe and adaptive long-term autonomy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(32)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.48550/arxiv.2211.09325
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
作者:
[Chuer Pan;Brian Okorn;Harry Zhang;Ben Eisner;David Held]
通讯作者:
Chuer Pan;Brian Okorn;Harry Zhang;Ben Eisner;David Held
Learning Off-policy for Online Planning
在线规划的离线学习
DOI:
--
发表时间:
2021
期刊:
Conference on Robot Learning (CoRL
影响因子:
--
作者:
[Sikchi, Harshit, Zhou, Wenxuan, Held, David]
通讯作者:
Held, David
DOI:
--
发表时间:
2022
期刊:
影响因子:
--
作者:
[Xingyu Lin;Carl Qi;Yunchu Zhang;Zhiao Huang;Katerina Fragkiadaki;Yunzhu Li;Chuang Gan;David Held]
通讯作者:
Xingyu Lin;Carl Qi;Yunchu Zhang;Zhiao Huang;Katerina Fragkiadaki;Yunzhu Li;Chuang Gan;David Held
DOI:
10.48550/arxiv.2211.11182
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
作者:
[Brian Okorn;Chuer Pan;M. Hebert;David Held]
通讯作者:
Brian Okorn;Chuer Pan;M. Hebert;David Held
DOI:
--
发表时间:
2021-11
期刊:
ArXiv
影响因子:
--
作者:
[Himangi Mittal;Brian Okorn;Arpit Jangid;David Held]
通讯作者:
Himangi Mittal;Brian Okorn;Arpit Jangid;David Held
共 27 条
CAREER: Self-supervised Representation Learning for Deformable Object Manipulation
-
批准号:2046491
-
项目类别:Continuing Grant
-
资助金额:$59.72万
-
财政年份:2021
-
负责人:David Held
-
依托单位:
国内基金
海外基金
Novosphingobium sp. FND-3降解呋喃丹的分子机制研究
-
批准号:31670112
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2016
-
负责人:洪青
-
依托单位: