The CoSTAR Block Stacking Dataset: Learning with Workspace Constraints

The CoSTAR Block Stacking Dataset: Learning with Workspace Constraints
复制标题

DOI:
10.1109/iros40897.2019.8967784
复制
发表时间:
2018-10
期刊:
2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Andrew Hundt;Varun Jain;Chia-Hung Lin;Chris Paxton;Gregory Hager
Andrew Hundt;Varun Jain;Chia-Hung Lin;Chris Paxton;Gregory Hager
中科院分区:
其他
文献类型:
--
作者:
Andrew Hundt;Varun Jain;Chia-Hung Lin;Chris Paxton;Gregory Hager

文献摘要

相似文献

机器人现在可以比以往任何时候都更有效地抓住物体,但一旦它拥有了物体,接下来会发生什么?我们发现,对现有对象抓取数据集中隐含的任务和工作空间约束的适度放松,可能会导致基于神经网络的抓取算法在更现实的环境下执行时,即使在简单的块堆叠任务上也会失败。为了解决这个问题,我们引入了JHU共星区块堆叠数据集(BSD),在该数据集中,机器人与5.1厘米彩色区块交互来完成订单履行式的区块堆叠任务。它包含动态场景和实时时间序列数据,与可比数据集相比,它在约束较少的环境中。有近12,000次堆叠尝试和200多万帧真实数据。我们将讨论该数据集如何为广泛的其他调查主题提供有价值的资源。我们发现,在以前的数据集上工作的手工设计的神经网络并不适用于这项任务。因此,为了建立该数据集的基准,我们使用一种新的多输入超树元模型来自动搜索基于神经网络的模型,并找到一个最终模型,该模型做出合理的3D姿势预测,以便在我们的数据集上抓取和堆叠。Costar BSD、代码和说明可在Sites.google.com/site/costardatet上找到。
A robot can now grasp an object more effectively than ever before, but once it has the object what happens next? We show that a mild relaxation of the task and workspace constraints implicit in existing object grasping datasets can cause neural network based grasping algorithms to fail on even a simple block stacking task when executed under more realistic circumstances. To address this, we introduce the JHU CoSTAR Block Stacking Dataset (BSD), where a robot interacts with 5.1 cm colored blocks to complete an order-fulfillment style block stacking task. It contains dynamic scenes and real time-series data in a less constrained environment than comparable datasets. There are nearly 12,000 stacking attempts and over 2 million frames of real data. We discuss the ways in which this dataset provides a valuable resource for a broad range of other topics of investigation. We find that hand-designed neural networks that work on prior datasets do not generalize to this task. Thus, to establish a baseline for this dataset, we demonstrate an automated search of neural network based models using a novel multiple-input HyperTree MetaModel, and find a final model which makes reasonable 3D pose predictions for grasping and stacking on our dataset. The CoSTAR BSD, code, and instructions are available at sites.google.com/site/costardataset.