Scene-level Pose Estimation for Multiple Instances of Densely Packed Objects

Scene-level Pose Estimation for Multiple Instances of Densely Packed Objects
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
--
影响因子:
--
通讯作者:
Chaitanya Mitash;Bowen Wen;Kostas E. Bekris;Abdeslam Boularias
Chaitanya Mitash;Bowen Wen;Kostas E. Bekris;Abdeslam Boularias
中科院分区:
其他
文献类型:
--
作者:
Chaitanya Mitash;Bowen Wen;Kostas E. Bekris;Abdeslam Boularias

文献摘要

被引文献

相似文献

本文介绍了关键的机器学习操作,这些操作允许从RGB-D数据中实现对密集堆积或非结构化堆积的对象的多个实例的稳健的、联合的6D姿态估计。第一个目标是学习语义和实例边界检测器,而不需要手动标记。对抗性训练框架与基于物理的模拟相结合被用来实现在合成数据和真实数据中行为相似的检测器。在给定这种检测器的随机输出的情况下,对对象姿势的候选进行采样。第二个目标是自动学习每个姿势候选的单个分数,该分数代表其通过渐变增强树解释整个场景的质量。该方法利用了从表面和边界对齐得到的特征,这些特征在被观测场景和以假设姿势放置的对象模型之间对齐。然后,通过整数线性规划过程来实现场景级的多实例姿态估计,该过程选择最大化所学习的个体分数的总和的假设,同时尊重约束,例如避免碰撞。为了评估该方法,收集了密集堆积对象的数据集,这些对象对于最先进的方法具有挑战性的设置。在这个数据集和一个公共数据集上的实验表明,该方法在仅使用合成数据集进行训练的情况下,在6D姿态精度方面明显优于其他方法。
This paper introduces key machine learning operations that allow the realization of robust, joint 6D pose estimation of multiple instances of objects either densely packed or in unstructured piles from RGB-D data. The first objective is to learn semantic and instance-boundary detectors without manual labeling. An adversarial training framework in conjunction with physics-based simulation is used to achieve detectors that behave similarly in synthetic and real data. Given the stochastic output of such detectors, candidates for object poses are sampled. The second objective is to automatically learn a single score for each pose candidate that represents its quality in terms of explaining the entire scene via a gradient boosted tree. The proposed method uses features derived from surface and boundary alignment between the observed scene and the object model placed at hypothesized poses. Scene-level, multi-instance pose estimation is then achieved by an integer linear programming process that selects hypotheses that maximize the sum of the learned individual scores, while respecting constraints, such as avoiding collisions. To evaluate this method, a dataset of densely packed objects with challenging setups for state-of-the-art approaches is collected. Experiments on this dataset and a public one show that the method significantly outperforms alternatives in terms of 6D pose accuracy while trained only with synthetic datasets.