TAMPC: A Controller for Escaping Traps in Novel Environments

TAMPC: A Controller for Escaping Traps in Novel Environments
复制标题

DOI:
10.1109/lra.2021.3057789
复制
发表时间:
2020-10
影响因子:
5.2
通讯作者:
Sheng Zhong;Zhenyuan Zhang;Nima Fazeli;D. Berenson
Sheng Zhong;Zhenyuan Zhang;Nima Fazeli;D. Berenson
中科院分区:
计算机科学2区
文献类型:
--
作者:
Sheng Zhong;Zhenyuan Zhang;Nima Fazeli;D. Berenson

文献摘要

被引文献

相似文献

我们提出了一种在混合和不连续动态的具有挑战性的情况下进行在线模型自适应和控制的方法,在给定控制器下,动作可能导致难以逃脱的“陷阱”状态。我们首先从随机收集的训练集中学习无陷阱系统的动态(因为我们不知道在线会遇到什么陷阱)。这些“标称”动态使我们能够在动态与训练数据匹配的场景中执行任务,但是当执行过程中出现意外陷阱时,我们必须找到一种方法来调整我们的动态和控制策略,并继续尝试任务。我们的方法,陷阱感知模型预测控制(TAMPC),是一种两级分层控制算法,它对陷阱和非标称动态进行推理,以在寻目标和恢复策略之间做出决策。我们方法的一个重要要求是,即使我们遇到相对于训练数据分布外的数据时,也能够识别标称动态。我们通过学习一种利用标称环境中的不变性的动态表示来实现这一点,从而实现更好的泛化。我们在模拟平面推动和轴孔装配以及真实机器人轴孔装配问题上,针对自适应控制、强化学习、陷阱处理基准方法评估了我们的方法,其中陷阱是由于我们只能通过接触观察到的意外障碍物而产生的。我们的结果表明,我们的方法在困难任务上优于基准方法,并且在较容易的任务上与先前的陷阱处理方法相当。
We propose an approach to online model adaptation and control in the challenging case of hybrid and discontinuous dynamics where actions may lead to difficult-to-escape “trap” states, under a given controller. We first learn dynamics for a system without traps from a randomly collected training set (since we do not know what traps will be encountered online). These “nominal” dynamics allow us to perform tasks in scenarios where the dynamics matches the training data, but when unexpected traps arise in execution, we must find a way to adapt our dynamics and control strategy and continue attempting the task. Our approach, Trap-Aware Model Predictive Control (TAMPC), is a two-level hierarchical control algorithm that reasons about traps and non-nominal dynamics to decide between goal-seeking and recovery policies. An important requirement of our method is the ability to recognize nominal dynamics even when we encounter data that is out-of-distribution w.r.t the training data. We achieve this by learning a representation for dynamics that exploits invariance in the nominal environment, thus allowing better generalization. We evaluate our method on simulated planar pushing and peg-in-hole as well as real robot peg-in-hole problems against adaptive control, reinforcement learning, trap-handling baselines, where traps arise due to unexpected obstacles that we only observe through contact. Our results show that our method outperforms the baselines on difficult tasks, and is comparable to prior trap-handling methods on easier tasks.