Certifiably-correct Control Policies for Safe Learning and Adaptation in Assistive Robotics

Certifiably-correct Control Policies for Safe Learning and Adaptation in Assistive Robotics
复制标题

DOI:
10.48550/arxiv.2303.06582
复制
发表时间:
2023-03
期刊:
ArXiv
影响因子:
--
通讯作者:
K. Majd;Geoffrey Clark;Tanmay Khandait;Siyu Zhou;S. Sankaranarayanan;Georgios Fainekos;H. B. Amor
K. Majd;Geoffrey Clark;Tanmay Khandait;Siyu Zhou;S. Sankaranarayanan;Georgios Fainekos;H. B. Amor
中科院分区:
其他
文献类型:
--
作者:
K. Majd;Geoffrey Clark;Tanmay Khandait;Siyu Zhou;S. Sankaranarayanan;Georgios Fainekos;H. B. Amor

文献摘要

相似文献

保证以人为中心的应用程序的安全性在机器人学习中至关重要,因为学习的策略可能会在以前看不见的场景中表现出不安全的行为。我们提出了一个框架,本地修复一个错误的策略网络,以满足一组正式的安全约束,使用混合二次规划(MIQP)。我们的MIQP公式明确地将安全约束施加到学习策略,同时最小化原始损失函数。然后验证策略网络在本地是安全的。我们展示了我们的框架的应用程序,以获得机器人小腿假肢的安全政策。
Guaranteeing safety in human-centric applications is critical in robot learning as the learned policies may demonstrate unsafe behaviors in formerly unseen scenarios. We present a framework to locally repair an erroneous policy network to satisfy a set of formal safety constraints using Mixed Integer Quadratic Programming (MIQP). Our MIQP formulation explicitly imposes the safety constraints to the learned policy while minimizing the original loss function. The policy network is then verified to be locally safe. We demonstrate the application of our framework to derive safe policies for a robotic lower-leg prosthesis.