课题基金 / 基金详情

SLES: SPECSRL: Specification-guided Perception-enabled Conformal Safe Reinforcement Learning

SLES: SPECSRL: Specification-guided Perception-enabled Conformal Safe Reinforcement Learning
SLES:SPECSRL:规范引导的感知启用的共形安全强化学习
批准号:
2331783
负责人:
Rajeev Alur
金额:
$150.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-10-01 至 2027-09-30

项目摘要

项目成果

Rajeev Alur的其他基金

相关文献

中文摘要
翻译
强化学习(RL)等机器学习技术有望成为自主家用机器人时代的关键支持技术,但为了实现这一承诺,必须将安全考虑纳入其设计中。除了避免对人类或机器人造成直接伤害外,家务劳动还会带来不同程度的额外危险,例如打翻热饮,忘记关水龙头,或压碎水果。SpecsRL项目提倡使用正式的高级逻辑规范来设计和部署安全的RL算法。这样的规范可以精确而简洁地传达机器人需要完成的预期目标以及期望避免的伤害。然后,可以设计用于策略综合的学习算法,并为综合策略满足规范的程度提供可量化的数学保证。为了实现规范引导RL的这一议程,并提供精确的数学和经验安全保证,SpecsRL汇集了具有强化学习,形式化方法,机器学习理论和机器人技术专业知识的研究人员。SpecsRL的主要贡献是为RL提供了一个新的框架,包括相关的任务规范语言、学习算法、保证和评估安全性的理论和经验技术,以及案例研究。SpecsRL的研究分为五个重点。(A)指定安全性:该项目开发了一种高度灵活的基于时间逻辑的规范语言,适用于具有可达性目标和软硬安全约束的机器人任务。一个关键的新奇之处在于,逻辑状态谓词(如“沸水”)是建立在视觉感知的基础上的。(B)共形安全策略综合:该项目开发了传播与可视化谓词和策略执行相关的基于共形预测的不确定性的方法,使整个系统能够派生端到端的正式安全保证。(C)安全恢复的在线干预:为了处理任务执行过程中具有高度不确定性的场景,该项目开发了在线监测和验证技术,允许机器人通过交互式行为“检查其工作”,例如串土豆以检查是否完成,从而允许稳健的决策和恢复。(D)经验安全测试:由于机器人可能会遇到在训练过程中没有考虑到的情况,该项目开发了经验压力测试技术来发现潜在的故障模式,以及适应分布变化的稳健方法。(E)实验评估:该方法在模拟厨房机器人和实验机器人厨房设施中进行了评估,以完成复杂的任务,如烹饪意大利面和设置餐桌。教育和知识转移活动是为了在RL和安全人工智能的交叉点建立一个研究人员社区,促进研究成果与教育的紧密结合,并促进多样性。这项研究得到了美国国家科学基金会和开放慈善机构的合作支持。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Machine learning techniques such as reinforcement learning (RL) promise to be the key enabling technology for an age of autonomous home robots, but to fulfill this promise, safety considerations must be integral to their design. In addition to avoiding immediate injury to humans or to the robot, household tasks carry additional hazards with varying levels of risk, such as spilling a hot drink, leaving a faucet on, or crushing a fruit. The SpecsRL project advocates the use of formal high-level logical specifications for design and deployment of safe RL algorithms. Such specifications can precisely and succinctly convey both the desired goals that a robot is required to accomplish as well as the harm it is expected to avoid.Learning algorithms for policy synthesis can then be designed with quantifiable mathematical guarantees for how well the synthesized policy meets the specification. To realize this agenda of specification-guided RL with precise mathematical and empirical safety guarantees, SpecsRL brings together researchers with expertise in reinforcement learning, formal methods, theory of machine learning, and robotics. The primary contribution of SpecsRL is a novel framework for RL with associated task specification language, learning algorithms, theoretical and empirical techniques for guaranteeing and evaluating safety, and case studies.SpecsRL research is organized along five thrusts. (A) Specifying safety: The project develops a highly flexible temporal-logic-based specification language suitable for robotic tasks with reachability goals and both hard and soft safety constraints. A key novelty is that logical state predicates (such as ``boiling water'') are grounded in visual perception. (B) Conformally safe policy synthesis: The project develops methods to propagate conformal prediction-based uncertainties associated with visual predicates and policy execution, enabling derivation of end-to-end formal safety guarantees for the overall system. (C) Online interventions for safe recovery: To deal with scenarios with high uncertainty during task execution, the project develops online monitoring and verification techniques that permit the robot to ``check its work'' through interactive behaviors, such as skewering a potato to check whether it is done, to permit robust decision-making and recovery. (D) Empirical safety testing: Since a robot is likely to encounter scenarios that are not considered during training, the project develops empirical stress testing techniques to discover potential failure modes, and robust ways to adjust to distribution shifts. (E) Experimental evaluation: The approach is evaluated for a kitchen robot in simulation as well as in experimental robot kitchen facility for complex tasks such as cooking pasta and setting a dinner table. The education and knowledge transfer activities are to build a community of researchers at the intersection of RL and safe AI, to facilitate tight integration of research results with education, and to promote diversity.This research is supported by a partnership between the National Science Foundation and Open Philanthropy.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
CCF: Medium: Enabling Real-Time Quantitative Decision Making over Streaming Data
  • 批准号:
    1763514
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $120.0万
  • 财政年份:
    2018
  • 负责人:
    Rajeev Alur
  • 依托单位:
SHF: Medium: Collaborative Research: Formal Analysis and Synthesis of Multiagent Systems with Incentives
  • 批准号:
    1703791
  • 项目类别:
    Standard Grant
  • 资助金额:
    $40.0万
  • 财政年份:
    2017
  • 负责人:
    Rajeev Alur
  • 依托单位:
Collaborative Research: Expeditions in Computer Augmented Program Engineering (ExCAPE): Harnessing Synthesis for Software Design
  • 批准号:
    1138996
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $375.0万
  • 财政年份:
    2012
  • 负责人:
    Rajeev Alur
  • 依托单位:
SHF: AF: SMALL: Scalable Symbolic Analysis of Hybrid Systems
  • 批准号:
    0915777
  • 项目类别:
    Standard Grant
  • 资助金额:
    $37.64万
  • 财政年份:
    2009
  • 负责人:
    Rajeev Alur
  • 依托单位: