DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with Tools

DiffSkill: Skill Abstraction from Differentiable Physics for Deformable Object Manipulations with Tools
复制标题

DOI:
10.48550/arxiv.2203.17275
复制
发表时间:
2022-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Xingyu Lin;Zhiao Huang;Yunzhu Li;J. Tenenbaum;David Held;Chuang Gan
Xingyu Lin;Zhiao Huang;Yunzhu Li;J. Tenenbaum;David Held;Chuang Gan
中科院分区:
其他
文献类型:
--
作者:
Xingyu Lin;Zhiao Huang;Yunzhu Li;J. Tenenbaum;David Held;Chuang Gan

文献摘要

被引文献

相似文献

我们考虑了机器人使用工具对可变形物体进行顺序操作的问题。先前的研究表明,可微物理模拟器为环境状态提供了梯度,并帮助轨迹优化以比无模型强化学习算法更快的速度收敛到可变形对象操作的数量级。然而,这种基于梯度的轨迹优化通常需要访问完整的模拟器状态,并且由于局部最优,只能解决短期的单技能任务。在这项工作中,我们提出了一个名为DiffSkill的新框架,该框架使用可微分物理模拟器进行技能抽象,以解决来自感官观察的长视界变形对象操作任务。特别是,我们首先使用基于梯度的优化器中的单个工具获得短视技能,在可微模拟器中使用完整的状态信息;然后,我们从以RGBD图像为输入的演示轨迹中学习神经技能抽象器。最后,我们通过寻找中间目标来规划技能,然后解决长期任务。与之前的强化学习算法和轨迹优化器相比,我们展示了我们的方法在一组新的顺序可变形对象操作任务中的优势。
We consider the problem of sequential robotic manipulation of deformable objects using tools. Previous works have shown that differentiable physics simulators provide gradients to the environment state and help trajectory optimization to converge orders of magnitude faster than model-free reinforcement learning algorithms for deformable object manipulation. However, such gradient-based trajectory optimization typically requires access to the full simulator states and can only solve short-horizon, single-skill tasks due to local optima. In this work, we propose a novel framework, named DiffSkill, that uses a differentiable physics simulator for skill abstraction to solve long-horizon deformable object manipulation tasks from sensory observations. In particular, we first obtain short-horizon skills using individual tools from a gradient-based optimizer, using the full state information in a differentiable simulator; we then learn a neural skill abstractor from the demonstration trajectories which takes RGBD images as input. Finally, we plan over the skills by finding the intermediate goals and then solve long-horizon tasks. We show the advantages of our method in a new set of sequential deformable object manipulation tasks compared to previous reinforcement learning algorithms and compared to the trajectory optimizer.