Learning Contraction Policies from Offline Data

Learning Contraction Policies from Offline Data
复制标题

DOI:
10.1109/lra.2022.3145100
复制
发表时间:
2021-12
影响因子:
5.2
通讯作者:
Navid Rezazadeh;Maxwell Kolarich;Solmaz S. Kia;Negar Mehr
Navid Rezazadeh;Maxwell Kolarich;Solmaz S. Kia;Negar Mehr
中科院分区:
计算机科学2区
文献类型:
--
作者:
Navid Rezazadeh;Maxwell Kolarich;Solmaz S. Kia;Negar Mehr

文献摘要

相似文献

本文提出了一种数据驱动的方法,学习收敛控制策略,从离线数据使用收缩理论。收缩理论使得能够构造使得闭环系统轨迹固有地收敛到唯一轨迹的策略。在技术层面,识别收缩度量(其是机器人的轨迹相对于其展现收缩的距离度量)通常是不平凡的。我们建议在实施收缩的同时共同学习控制策略及其相应的收缩度量。为了实现这一点,我们学习的机器人系统的隐式动力学模型,从离线数据集组成的机器人的状态和输入轨迹。使用这个学习的动力学模型,我们提出了一个数据增强算法学习收缩政策。我们在状态空间中随机生成样本,并通过学习的动力学模型将它们在时间上向前传播,以生成辅助样本轨迹。然后,我们学习控制策略和收缩度量,使得来自离线数据集的轨迹与我们生成的辅助样本轨迹之间的距离随着时间的推移而减小。我们评估了我们提出的框架在模拟机器人目标达成任务上的性能,并证明了强制收缩会导致更快的收敛速度和更强的鲁棒性。
This paper proposes a data-driven method for learning convergent control policies from offline data using Contraction theory. Contraction theory enables constructing a policy that makes the closed-loop system trajectories inherently convergent towards a unique trajectory. At the technical level, identifying the contraction metric, which is the distance metric with respect to which a robot's trajectories exhibit contraction is often non-trivial. We propose to jointly learn the control policy and its corresponding contraction metric while enforcing contraction. To achieve this, we learn an implicit dynamics model of the robotic system from an offline data set consisting of the robot's state and input trajectories. Using this learned dynamics model, we propose a data augmentation algorithm for learning contraction policies. We randomly generate samples in the state-space and propagate them forward in time through the learned dynamics model to generate auxiliary sample trajectories. We then learn both the control policy and the contraction metric such that the distance between the trajectories from the offline data set and our generated auxiliary sample trajectories decreases over time. We evaluate the performance of our proposed framework on simulated robotic goal-reaching tasks and demonstrate that enforcing contraction results in faster convergence and greater robustness of the learned policy.