SLES: SPECSRL: Specification-guided Perception-enabled Conformal Safe Reinforcement Learning
SLES: SPECSRL: Specification-guided Perception-enabled Conformal Safe Reinforcement Learning
批准号:
2331783
负责人:
Rajeev Alur
金额:
$150.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2027-09-30
中文摘要
强化学习(RL)等机器学习技术有望成为自主家庭机器人时代的关键使能技术,但要实现这一承诺,安全考虑必须是它们设计中不可或缺的一部分。除了避免对人类或机器人造成直接伤害外,家务劳动还会带来额外的危险,具有不同程度的风险,如洒出热饮、开着水龙头或碾碎水果。SpesRL项目提倡使用正式的高级逻辑规范来设计和部署安全的RL算法。这样的规范可以准确而简洁地表达机器人需要完成的期望目标以及期望避免的危害。然后,可以设计用于策略合成的学习算法,以量化数学保证合成的策略满足规范的程度。为了实现规范指导的RL议程,并提供精确的数学和经验安全保证,SpesRL汇集了具有强化学习、形式化方法、机器学习理论和机器人学专业知识的研究人员。SpecsRL的主要贡献是为RL提供了一个新的框架,包括相关的任务描述语言、学习算法、用于保证和评估安全性的理论和经验技术以及案例研究。(A)指定安全性:该项目开发了一种高度灵活的基于时态逻辑的规范语言,适用于具有可达性目标以及硬安全和软安全约束的机器人任务。一个关键的新奇之处在于,逻辑状态谓词(如“沸水”)是基于视觉感知的。(B)共形安全政策合成:该项目开发了传播与可视谓词和政策执行相关的基于共形预测的不确定性的方法,从而能够推导出整个系统的端到端正式安全保证。(C)安全恢复的在线干预:为了处理任务执行过程中具有高度不确定性的情况,该项目开发了在线监测和核查技术,使机器人能够通过互动行为“检查其工作”,例如将土豆串起来检查是否完成,从而作出强有力的决策和恢复。(D)经验安全测试:由于机器人可能会遇到培训期间未考虑到的情况,该项目开发了经验压力测试技术,以发现潜在的故障模式,并采用可靠的方法来适应分布变化。(E)实验评估:对厨房机器人在模拟中以及在实验机器人厨房设施中执行复杂任务(如烹调意大利面和摆放餐桌)进行了评估。教育和知识转移活动是为了在RL和SAFE AI的交叉点建立一个研究人员社区,促进研究成果与教育的紧密结合,并促进多样性。这项研究得到了国家科学基金会和开放慈善机构的合作伙伴关系的支持。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine learning techniques such as reinforcement learning (RL) promise to be the key enabling technology for an age of autonomous home robots, but to fulfill this promise, safety considerations must be integral to their design. In addition to avoiding immediate injury to humans or to the robot, household tasks carry additional hazards with varying levels of risk, such as spilling a hot drink, leaving a faucet on, or crushing a fruit. The SpecsRL project advocates the use of formal high-level logical specifications for design and deployment of safe RL algorithms. Such specifications can precisely and succinctly convey both the desired goals that a robot is required to accomplish as well as the harm it is expected to avoid.Learning algorithms for policy synthesis can then be designed with quantifiable mathematical guarantees for how well the synthesized policy meets the specification. To realize this agenda of specification-guided RL with precise mathematical and empirical safety guarantees, SpecsRL brings together researchers with expertise in reinforcement learning, formal methods, theory of machine learning, and robotics. The primary contribution of SpecsRL is a novel framework for RL with associated task specification language, learning algorithms, theoretical and empirical techniques for guaranteeing and evaluating safety, and case studies.SpecsRL research is organized along five thrusts. (A) Specifying safety: The project develops a highly flexible temporal-logic-based specification language suitable for robotic tasks with reachability goals and both hard and soft safety constraints. A key novelty is that logical state predicates (such as ``boiling water'') are grounded in visual perception. (B) Conformally safe policy synthesis: The project develops methods to propagate conformal prediction-based uncertainties associated with visual predicates and policy execution, enabling derivation of end-to-end formal safety guarantees for the overall system. (C) Online interventions for safe recovery: To deal with scenarios with high uncertainty during task execution, the project develops online monitoring and verification techniques that permit the robot to ``check its work'' through interactive behaviors, such as skewering a potato to check whether it is done, to permit robust decision-making and recovery. (D) Empirical safety testing: Since a robot is likely to encounter scenarios that are not considered during training, the project develops empirical stress testing techniques to discover potential failure modes, and robust ways to adjust to distribution shifts. (E) Experimental evaluation: The approach is evaluated for a kitchen robot in simulation as well as in experimental robot kitchen facility for complex tasks such as cooking pasta and setting a dinner table. The education and knowledge transfer activities are to build a community of researchers at the intersection of RL and safe AI, to facilitate tight integration of research results with education, and to promote diversity.This research is supported by a partnership between the National Science Foundation and Open Philanthropy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CCF: Medium: Enabling Real-Time Quantitative Decision Making over Streaming Data
-
批准号:1763514
-
项目类别:Continuing Grant
-
资助金额:$120.0万
-
财政年份:2018
-
负责人:Rajeev Alur
-
依托单位:
SHF: Medium: Collaborative Research: Formal Analysis and Synthesis of Multiagent Systems with Incentives
-
批准号:1703791
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2017
-
负责人:Rajeev Alur
-
依托单位:
Collaborative Research: Expeditions in Computer Augmented Program Engineering (ExCAPE): Harnessing Synthesis for Software Design
-
批准号:1138996
-
项目类别:Continuing Grant
-
资助金额:$375.0万
-
财政年份:2012
-
负责人:Rajeev Alur
-
依托单位:
SHF: AF: SMALL: Scalable Symbolic Analysis of Hybrid Systems
-
批准号:0915777
-
项目类别:Standard Grant
-
资助金额:$37.64万
-
财政年份:2009
-
负责人:Rajeev Alur
-
依托单位:
SHF: Medium: Formal Analysis of Concurrent Software on Relaxed Memory Models
-
批准号:0905464
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2009
-
负责人:Rajeev Alur
-
依托单位:
Behavioral Interfaces for Software Components
-
批准号:0541149
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2006
-
负责人:Rajeev Alur
-
依托单位:
Proposal for Hybrid Systems Workshop; March 25-28, 2004, Philadelphia, PA
-
批准号:0401049
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:2004
-
负责人:Rajeev Alur
-
依托单位:
Synthesis of Embedded Software from Hybrid Models
-
批准号:0410662
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2004
-
负责人:Rajeev Alur
-
依托单位:
WORKSHOP ON EMBEDDED SOFTWARE
-
批准号:0318299
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2003
-
负责人:Rajeev Alur
-
依托单位:
GAMES FOR FORMAL DESIGN AND VERIFICATION OF REACTIVE SYSTEMS
-
批准号:0306382
-
项目类别:Standard Grant
-
资助金额:$27.0万
-
财政年份:2003
-
负责人:Rajeev Alur
-
依托单位:
ITR/SY: Formal Design and Analysis of Hybrid Systems
-
批准号:0121431
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2001
-
负责人:Rajeev Alur
-
依托单位:
Specification, Analysis, and Testing of Scenario-Based Requirements
-
批准号:9970925
-
项目类别:Continuing Grant
-
资助金额:$21.5万
-
财政年份:1999
-
负责人:Rajeev Alur
-
依托单位:
CAREER: Computer-Aided Verification of Reactive Systems
-
批准号:9734115
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:1998
-
负责人:Rajeev Alur
-
依托单位: