Collaborative Research: SLES: Safe Distributional-Reinforcement Learning-Enabled Systems: Theories, Algorithms, and Experiments
Collaborative Research: SLES: Safe Distributional-Reinforcement Learning-Enabled Systems: Theories, Algorithms, and Experiments
批准号:
2331781
负责人:
Wenlong Zhang
金额:
$75.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2027-09-30
中文摘要
强化学习(RL)因其在自动化和机器人学方面的成功而被广泛视为下一代学习系统最重要的技术之一。例如,6G网络系统、自动驾驶、数字医疗和智慧城市都是由RL实现的。然而,尽管在过去的几十年里取得了显著的进步,但在实践中应用RL的一个主要障碍是缺乏健壮性、对尾部风险的弹性、操作约束等“安全”保证。这是因为传统的RL只以最大化累积回报为目标。虽然在传统的RL算法中可以将惩罚添加到奖励中以阻止不安全的行为,但许多安全约束,如机会约束,不能简单地视为惩罚。该项目开发了基于分布式强化学习(DRL)的安全RL使能系统的基础技术,它学习最优策略。在为支持安全学习的系统开发DRL基础的同时,通过将该项目中开发的新理论和算法纳入研究生课程,研究和教育被整合在一起。所有团队成员一直在定期监督本科生和来自代表性不足群体的学生。该团队继续利用俄亥俄州立大学的女性地位和亚利桑那州立大学的女性科学与工程项目,以加强女学生和研究人员的更广泛参与。该项目侧重于支持DRL的系统的端到端安全的综合方法。端到端安全包括:(1)政策安全:学习避免灾难性后果发生的安全政策(对应于对风险敏感的RL);(2)勘探安全--通过避免勘探/学习期间的危险行动安全地学习安全政策(对应于在线RL);以及(3)环境安全--学习对参数不确定性(环境变化)稳健的政策。该项目包括四个推进器。约束DRL的基础1旨在建立风险敏感的约束DRL的理论基础,重点关注政策和环境安全。推力2(在线受限DRL)考虑安全的在线学习和决策,并在学习安全的DRL政策时专注于勘探安全和环境安全。推力3(物理增强的受限DRL)利用物理来增强端到端的安全性。这三项基础性研究是相互依存的,但每一项都侧重于支持安全RL的系统的一个独特方面,并解决了多种安全概念。第四个推力将通过高保真模拟和使用无人机的真实世界实验提供全面的验证。这项研究由国家科学基金会和开放慈善机构之间的合作伙伴关系支持。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Reinforcement learning (RL), with its success in automation and robotics, has been widely viewed as one of the most important technologies for next-generation, learning-enabled systems. For example, 6G networking systems, autonomous driving, digital healthcare, and smart cities are all enabled by RL. However, despite the significant advances over the last few decades, a major obstacle in applying RL in practice is the lack of “safety'' guarantees such as robustness, resilience to tail-risks, operational constraints, etc. This is because the traditional RL only aims at maximizing cumulative reward. While it is possible to add penalties to rewards in a traditional RL algorithm to discourage unsafe actions, many safety constraints, such as chance constraints, cannot be simply treated as penalties. This project develops foundational technologies for safe RL-enabled systems based on Distributional Reinforcement Learning (DRL), which learns the optimal policy. While developing the foundation of DRL for safe learning-enabled systems, research and education are integrated by including new theories and algorithms developed in this project into their graduate-level courses. All team members have been regularly supervising undergraduate students and students from underrepresented groups. The team continues to leverage Women's Place at Ohio State University and the Women in Science and Engineering Program at Arizona State University to enhance the broader participation of women students and researchers. This project focuses on a comprehensive approach for the end-to-end safety of DRL-enabled systems. The end-to-end safety includes (i) policy safety: learn a safe policy to avoid the occurrence of catastrophic outcomes (corresponds to risk-sensitive RL); (ii) exploration safety -- learn a safe policy safely by avoiding dangerous actions during exploration/learning (corresponds to online RL); and (iii) environmental safety -- learn a policy that is robust to parametric uncertainty (environment change). This project includes four thrusts. Thrust 1 (Foundation of constrained DRL) aims to establish theoretical foundations of risk sensitive constrained DRL and focuses on policy and environmental safety. Thrust 2 (Online constrained DRL) considers safe online learning and decision-making and focuses on exploration safety and environmental safety when learning a safe DRL policy. Thrust 3 (Physics-Enhanced constrained DRL) exploits physics to enhance end-to-end safety. These three thrusts on foundational research are interdependent, but each focuses on a unique aspect of safe RL-enabled systems and addresses multiple safety notions. The fourth thrust will provide comprehensive validation with both high-fidelity simulations and real-world experiments using unmanned aerial vehicles.This research is supported by a partnership between the National Science Foundation and Open Philanthropy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CCRI: Planning-C: Developing a Minecraft-based Testbed for Evaluating Human-AI Teaming Research
-
批准号:2213827
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2022
-
负责人:Wenlong Zhang
-
依托单位:
I-Corps: Wearable Soft Robotic Glove for Hand Assistance and Rehabilitation
-
批准号:2132714
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2021
-
负责人:Wenlong Zhang
-
依托单位:
CAREER: Facilitating Human Interaction with Assistive Robots Through Intent Signaling and Inference
-
批准号:1944833
-
项目类别:Standard Grant
-
资助金额:$55.18万
-
财政年份:2020
-
负责人:Wenlong Zhang
-
依托单位:
NRI: FND: Scalable and Customizable Intent Inference and Motion Planning for Socially-Adept Autonomous Vehicles
-
批准号:1925403
-
项目类别:Standard Grant
-
资助金额:$75.0万
-
财政年份:2019
-
负责人:Wenlong Zhang
-
依托单位:
CRII: CHS: Enabling Safe and Adaptive Robot-aided Gait Training through Biomechanical Characterization and Learning from Demonstration
-
批准号:1756031
-
项目类别:Standard Grant
-
资助金额:$17.5万
-
财政年份:2018
-
负责人:Wenlong Zhang
-
依托单位:
EAGER: Distributed Iterative Control of Soft Robotic Arms
-
批准号:1800940
-
项目类别:Standard Grant
-
资助金额:$14.31万
-
财政年份:2018
-
负责人:Wenlong Zhang
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Research on Quantum Field Theory without a Lagrangian Description
-
批准号:24ZR1403900
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:SATOSHI NAWATA
-
依托单位:
Cell Research
-
批准号:31224802
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2012
-
负责人:程磊
-
依托单位:
Cell Research
-
批准号:31024804
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2010
-
负责人:程磊
-
依托单位:
Cell Research (细胞研究)
-
批准号:30824808
-
项目类别:专项基金项目
-
资助金额:24.0万元
-
批准年份:2008
-
负责人:张爱兰
-
依托单位:
Research on the Rapid Growth Mechanism of KDP Crystal
-
批准号:10774081
-
项目类别:面上项目
-
资助金额:45.0万元
-
批准年份:2007
-
负责人:滕冰
-
依托单位: