SLES: SPECSRL: Specification-guided Perception-enabled Conformal Safe Reinforcement Learning
SLES: SPECSRL: Specification-guided Perception-enabled Conformal Safe Reinforcement Learning
批准号:
2331783
负责人:
Rajeev Alur
金额:
$150.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2027-09-30
中文摘要
强化学习(RL)等机器学习技术有望成为自主家用机器人时代的关键技术,但为了实现这一承诺,安全考虑必须成为其设计的一部分。除了避免对人类或机器人造成直接伤害外,家务劳动还带有不同程度的危险,例如打翻热饮,打开水龙头或压碎水果。SpecsRL项目提倡使用正式的高级逻辑规范来设计和部署安全的RL算法。这样的规范可以精确而简洁地传达机器人需要完成的预期目标以及期望避免的危害。然后,可以设计用于策略合成的学习算法,并对合成策略满足规范的程度进行量化的数学保证。为了实现这一具有精确数学和经验安全保证的规范指导RL议程,SpecsRL汇集了在强化学习,形式化方法,机器学习理论和机器人技术方面具有专业知识的研究人员。SpecsRL的主要贡献是为RL提供了一个新的框架,包括相关的任务规范语言、学习算法、用于保证和评估安全性的理论和经验技术以及案例研究。SpecsRL的研究沿着沿着五个方向组织。(A)安全性:该项目开发了一种高度灵活的基于时序逻辑的规范语言,适用于具有可达性目标和硬、软安全约束的机器人任务。一个关键的新奇在于逻辑状态谓词(如“沸水”)是基于视觉感知的。(B)共形安全策略合成:该项目开发了传播与视觉谓词和策略执行相关的基于共形预测的不确定性的方法,从而为整个系统提供端到端的正式安全保证。(C)安全恢复的在线干预:为了处理任务执行过程中高度不确定的情况,该项目开发了在线监测和验证技术,允许机器人通过互动行为“检查其工作”,例如串土豆以检查是否完成,以允许强大的决策和恢复。(D)经验性安全性测试:由于机器人很可能会遇到在训练过程中没有考虑到的场景,该项目开发了经验压力测试技术来发现潜在的故障模式,以及适应分布变化的鲁棒方法。(E)实验评价:该方法进行评估的厨房机器人在模拟以及在实验机器人厨房设施复杂的任务,如烹饪意大利面和设置餐桌。教育和知识转移活动是在强化学习和安全人工智能的交叉点上建立一个研究人员社区,以促进研究成果与教育的紧密结合,这项研究得到了美国国家科学基金会和开放慈善机构之间的合作伙伴关系的支持。该奖项反映了NSF的法定使命,并通过使用基金会的知识产权进行评估,被认为值得支持。优点和更广泛的影响审查标准。
英文摘要
Machine learning techniques such as reinforcement learning (RL) promise to be the key enabling technology for an age of autonomous home robots, but to fulfill this promise, safety considerations must be integral to their design. In addition to avoiding immediate injury to humans or to the robot, household tasks carry additional hazards with varying levels of risk, such as spilling a hot drink, leaving a faucet on, or crushing a fruit. The SpecsRL project advocates the use of formal high-level logical specifications for design and deployment of safe RL algorithms. Such specifications can precisely and succinctly convey both the desired goals that a robot is required to accomplish as well as the harm it is expected to avoid.Learning algorithms for policy synthesis can then be designed with quantifiable mathematical guarantees for how well the synthesized policy meets the specification. To realize this agenda of specification-guided RL with precise mathematical and empirical safety guarantees, SpecsRL brings together researchers with expertise in reinforcement learning, formal methods, theory of machine learning, and robotics. The primary contribution of SpecsRL is a novel framework for RL with associated task specification language, learning algorithms, theoretical and empirical techniques for guaranteeing and evaluating safety, and case studies.SpecsRL research is organized along five thrusts. (A) Specifying safety: The project develops a highly flexible temporal-logic-based specification language suitable for robotic tasks with reachability goals and both hard and soft safety constraints. A key novelty is that logical state predicates (such as ``boiling water'') are grounded in visual perception. (B) Conformally safe policy synthesis: The project develops methods to propagate conformal prediction-based uncertainties associated with visual predicates and policy execution, enabling derivation of end-to-end formal safety guarantees for the overall system. (C) Online interventions for safe recovery: To deal with scenarios with high uncertainty during task execution, the project develops online monitoring and verification techniques that permit the robot to ``check its work'' through interactive behaviors, such as skewering a potato to check whether it is done, to permit robust decision-making and recovery. (D) Empirical safety testing: Since a robot is likely to encounter scenarios that are not considered during training, the project develops empirical stress testing techniques to discover potential failure modes, and robust ways to adjust to distribution shifts. (E) Experimental evaluation: The approach is evaluated for a kitchen robot in simulation as well as in experimental robot kitchen facility for complex tasks such as cooking pasta and setting a dinner table. The education and knowledge transfer activities are to build a community of researchers at the intersection of RL and safe AI, to facilitate tight integration of research results with education, and to promote diversity.This research is supported by a partnership between the National Science Foundation and Open Philanthropy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CCF: Medium: Enabling Real-Time Quantitative Decision Making over Streaming Data
-
批准号:1763514
-
项目类别:Continuing Grant
-
资助金额:$120.0万
-
财政年份:2018
-
负责人:Rajeev Alur
-
依托单位:
SHF: Medium: Collaborative Research: Formal Analysis and Synthesis of Multiagent Systems with Incentives
-
批准号:1703791
-
项目类别:Standard Grant
-
资助金额:$40.0万
-
财政年份:2017
-
负责人:Rajeev Alur
-
依托单位:
Collaborative Research: Expeditions in Computer Augmented Program Engineering (ExCAPE): Harnessing Synthesis for Software Design
-
批准号:1138996
-
项目类别:Continuing Grant
-
资助金额:$375.0万
-
财政年份:2012
-
负责人:Rajeev Alur
-
依托单位:
SHF: AF: SMALL: Scalable Symbolic Analysis of Hybrid Systems
-
批准号:0915777
-
项目类别:Standard Grant
-
资助金额:$37.64万
-
财政年份:2009
-
负责人:Rajeev Alur
-
依托单位:
SHF: Medium: Formal Analysis of Concurrent Software on Relaxed Memory Models
-
批准号:0905464
-
项目类别:Standard Grant
-
资助金额:$120.0万
-
财政年份:2009
-
负责人:Rajeev Alur
-
依托单位:
Behavioral Interfaces for Software Components
-
批准号:0541149
-
项目类别:Standard Grant
-
资助金额:$30.0万
-
财政年份:2006
-
负责人:Rajeev Alur
-
依托单位:
Proposal for Hybrid Systems Workshop; March 25-28, 2004, Philadelphia, PA
-
批准号:0401049
-
项目类别:Standard Grant
-
资助金额:$2.0万
-
财政年份:2004
-
负责人:Rajeev Alur
-
依托单位:
Synthesis of Embedded Software from Hybrid Models
-
批准号:0410662
-
项目类别:Continuing Grant
-
资助金额:$40.0万
-
财政年份:2004
-
负责人:Rajeev Alur
-
依托单位:
WORKSHOP ON EMBEDDED SOFTWARE
-
批准号:0318299
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2003
-
负责人:Rajeev Alur
-
依托单位:
GAMES FOR FORMAL DESIGN AND VERIFICATION OF REACTIVE SYSTEMS
-
批准号:0306382
-
项目类别:Standard Grant
-
资助金额:$27.0万
-
财政年份:2003
-
负责人:Rajeev Alur
-
依托单位:
ITR/SY: Formal Design and Analysis of Hybrid Systems
-
批准号:0121431
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2001
-
负责人:Rajeev Alur
-
依托单位:
Specification, Analysis, and Testing of Scenario-Based Requirements
-
批准号:9970925
-
项目类别:Continuing Grant
-
资助金额:$21.5万
-
财政年份:1999
-
负责人:Rajeev Alur
-
依托单位:
CAREER: Computer-Aided Verification of Reactive Systems
-
批准号:9734115
-
项目类别:Continuing Grant
-
资助金额:$20.0万
-
财政年份:1998
-
负责人:Rajeev Alur
-
依托单位: