On-Policy Dataset Synthesis for Learning Robot Grasping Policies Using Fully Convolutional Deep Networks

On-Policy Dataset Synthesis for Learning Robot Grasping Policies Using Fully Convolutional Deep Networks
复制标题

DOI:
10.1109/lra.2019.2895878
复制
发表时间:
2019-04-01
影响因子:
5.2
通讯作者:
Goldberg, Ken
Goldberg, Ken
中科院分区:
计算机科学2区
文献类型:
--
作者:
Satish, Vishal;Mahler, Jeffrey;Goldberg, Ken

文献摘要

被引文献

相似文献

快速可靠的机器人抓住各种对象的应用程序,具有从仓库自动化到家庭整洁的应用。一种有希望的方法是从点云的合成训练数据集,grasps和奖励的综合训练数据集学习,并使用带有随机噪声模型进行域随机化的分析模型来采样。在这封信中,我们探讨了合成培训示例的分布如何影响学到的机器人政策的速度和可靠性。我们提出了一个合成数据采样分布,该分布结合了从策略动作设置的grasps和具有完整状态知识的强大掌握主管中的指导样本。我们使用它来培训基于完全卷积网络体系结构的机器人政策,该架构评估了数百万个以4-DOF(3-D位置和平面定向)的求解。物理机器人实验表明,基于完全卷积的掌握质量CNN(FC-GQ-CNN)的策略可以计划在0.625 s的范围内,考虑到基于迭代的抓取样本和评估的先前政策比我们先前的策略多5000倍。该计算效率提高了速率和可靠性,每小时达到296次平均选择(MPPH),而迭代策略为250 mpph。灵敏度实验探讨了主管指导水平和政策行动空间颗粒状的影响。可以在http://berkeleyautomation.github.io/fcgqccnn上找到代码,数据集,视频和补充材料。
Rapid and reliable robot grasping for a diverse set of objects has applications from warehouse automation to home de-cluttering. One promising approach is to learn deep policies from synthetic training datasets of point clouds, grasps, and rewards sampled using analytic models with stochastic noise models for domain randomization. In this letter, we explore how the distribution of synthetic training examples affects the rate and reliability of the learned robot policy. We propose a synthetic data sampling distribution that combines grasps sampled from the policy action set with guiding samples from a robust grasping supervisor that has full state knowledge. We use this to train a robot policy based on a fully convolutional network architecture that evaluates millions of grasp candidates in 4-DOF (3-D position and planar orientation). Physical robot experiments suggest that a policy based on fully convolutional grasp quality CNNs (FC-GQ-CNNs) can plan grasps in 0.625 s, considering 5000x more grasps than our prior policy based on iterative grasp sampling and evaluation. This computational efficiency improves rate and reliability, achieving 296 mean picks per hour (MPPH) compared to 250 MPPH for iterative policies. Sensitivity experiments explore the effect of supervisor guidance level and granularity of the policy action space. Code, datasets, videos, and supplementary material can be found at http://berkeleyautomation.github.io/fcgqcnn.