Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects

Deep Object Pose Estimation for Semantic Robotic Grasping of Household Objects
复制标题

DOI:
--
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Jonathan Tremblay;Thang To;Balakumar Sundaralingam;Yu Xiang;D. Fox;Stan Birchfield
Jonathan Tremblay;Thang To;Balakumar Sundaralingam;Yu Xiang;D. Fox;Stan Birchfield
中科院分区:
其他
文献类型:
--
作者:
Jonathan Tremblay;Thang To;Balakumar Sundaralingam;Yu Xiang;D. Fox;Stan Birchfield

文献摘要

被引文献

相似文献

使用合成数据来训练机器人操作的深度神经网络,有望获得几乎无限量的预先标记的训练数据,这些数据是安全生成的。迄今为止,合成数据的关键挑战之一是弥合所谓的现实差距,以便在合成数据上训练的网络在暴露于真实世界数据时能够正确运行。我们探索现实差距的背景下,已知的物体从一个单一的RGB图像的6自由度姿态估计。我们表明,对于这个问题的现实差距可以成功地跨越域随机化和逼真的数据的简单组合。使用以这种方式生成的合成数据,我们引入了一个一次性深度神经网络,该网络能够与根据真实的和合成数据组合训练的最先进网络进行竞争。据我们所知,这是第一个仅在合成数据上训练的深度网络,能够在6自由度物体姿态估计上实现最先进的性能。我们的网络还可以更好地推广到新的环境,包括极端的照明条件,我们显示了定性的结果。使用这个网络,我们展示了一个实时系统,估计对象的姿态与足够的准确性,为现实世界的语义掌握已知的家居对象的杂乱的真实的机器人。
Using synthetic data for training deep neural networks for robotic manipulation holds the promise of an almost unlimited amount of pre-labeled training data, generated safely out of harm's way. One of the key challenges of synthetic data, to date, has been to bridge the so-called reality gap, so that networks trained on synthetic data operate correctly when exposed to real-world data. We explore the reality gap in the context of 6-DoF pose estimation of known objects from a single RGB image. We show that for this problem the reality gap can be successfully spanned by a simple combination of domain randomized and photorealistic data. Using synthetic data generated in this manner, we introduce a one-shot deep neural network that is able to perform competitively against a state-of-the-art network trained on a combination of real and synthetic data. To our knowledge, this is the first deep network trained only on synthetic data that is able to achieve state-of-the-art performance on 6-DoF object pose estimation. Our network also generalizes better to novel environments including extreme lighting conditions, for which we show qualitative results. Using this network we demonstrate a real-time system estimating object poses with sufficient accuracy for real-world semantic grasping of known household objects in clutter by a real robot.