Signal Temporal Logic Neural Predictive Control

Signal Temporal Logic Neural Predictive Control
复制标题

DOI:
10.1109/lra.2023.3315536
复制
发表时间:
2023-09
影响因子:
5.2
通讯作者:
Yue Meng;Chuchu Fan
Yue Meng;Chuchu Fan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yue Meng;Chuchu Fan

文献摘要

相似文献

确保安全并满足时间规范是长期机器人任务的关键挑战。信号时序逻辑(STL)已被广泛用于系统地、严格地指定这些要求。然而,在这些 STL 要求下寻找控制策略的传统方法计算复杂,并且无法扩展到高维或具有复杂非线性动力学的系统。强化学习 (RL) 方法可以通过手工设计或受 STL 启发的奖励来学习满足 STL 规范的策略,但可能会由于奖励的模糊性和稀疏性而遇到意外行为。在这封信中,我们提出了一种直接学习神经网络控制器的方法,以满足 STL 中指定的要求。我们的控制器学习推出轨迹,以在训练中最大化 STL 鲁棒性得分。在测试中,与模型预测控制 (MPC) 类似,学习控制器预测规划范围内的轨迹,以确保满足部署中的 STL 要求。备份策略旨在确保控制器发生故障时的安全。我们的方法可以适应各种初始条件和环境参数。我们对六个任务进行了实验,其中我们的备份策略方法在 STL 满足率方面优于经典方法(MPC、STL 求解器)、无模型和基于模型的 RL 方法,特别是在具有复杂 STL 规范的任务上,同时比经典方法快 10 倍到 100 倍。
Ensuring safety and meeting temporal specifications are critical challenges for long-term robotic tasks. Signal temporal logic (STL) has been widely used to systematically and rigorously specify these requirements. However, traditional methods of finding the control policy under those STL requirements are computationally complex and not scalable to high-dimensional or systems with complex nonlinear dynamics. Reinforcement learning (RL) methods can learn the policy to satisfy the STL specifications via hand-crafted or STL-inspired rewards, but might encounter unexpected behaviors due to ambiguity and sparsity in the reward. In this letter, we propose a method to directly learn a neural network controller to satisfy the requirements specified in STL. Our controller learns to roll out trajectories to maximize the STL robustness score in training. In testing, similar to Model Predictive Control (MPC), the learned controller predicts a trajectory within a planning horizon to ensure the satisfaction of the STL requirement in deployment. A backup policy is designed to ensure safety when our controller fails. Our approach can adapt to various initial conditions and environmental parameters. We conduct experiments on six tasks, where our method with the backup policy outperforms the classical methods (MPC, STL-solver), model-free and model-based RL methods in STL satisfaction rate, especially on tasks with complex STL specifications while being 10X-100X faster than the classical methods.