Visual Identification of Articulated Object Parts

Visual Identification of Articulated Object Parts
复制标题

DOI:
10.1109/iros51168.2021.9636054
复制
发表时间:
2020-12
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
V. Zeng;Timothy E. Lee;Jacky Liang;Oliver Kroemer
V. Zeng;Timothy E. Lee;Jacky Liang;Oliver Kroemer
中科院分区:
其他
文献类型:
--
作者:
V. Zeng;Timothy E. Lee;Jacky Liang;Oliver Kroemer

文献摘要

被引文献

相似文献

当自主机器人在诸如家庭的真实世界环境中进行交互和导航时,可靠地识别和操纵铰接对象(例如门和橱柜)是有用的。在物体关节识别中的许多现有工作需要由机器人或人类操纵物体。虽然最近的工作已经解决了预测关节类型从视觉观察单独,他们往往假设先验知识的类别级运动学运动模型或序列的观察,其中的关节部分移动根据其运动学约束。在这项工作中,我们提出了FormNet,这是一种神经网络,可以识别RGB-D图像和分割掩模的单帧中对象部分对之间的清晰度机制。该网络在来自6个类别的149个铰接对象的10万张合成图像上进行训练。合成图像通过具有域随机化的真实感模拟器渲染。我们提出的模型预测运动残留流的对象部分,这些流量被用来确定关节类型和参数。该网络在训练类别中的新对象实例上实现了82.5%的清晰度类型分类准确率。实验还表明,该方法可以推广到新的类别,并可以应用于现实世界的图像,而无需微调。
As autonomous robots interact and navigate around real-world environments such as homes, it is useful to reliably identify and manipulate articulated objects, such as doors and cabinets. Many prior works in object articulation identification require manipulation of the object, either by the robot or a human. While recent works have addressed predicting articulation types from visual observations alone, they often assume prior knowledge of category-level kinematic motion models or sequence of observations where the articulated parts are moving according to their kinematic constraints. In this work, we propose FormNet, a neural network that identifies the articulation mechanisms between pairs of object parts from a single frame of an RGB-D image and segmentation masks. The network is trained on 100k synthetic images of 149 articulated objects from 6 categories. Synthetic images are rendered via a photorealistic simulator with domain randomization. Our proposed model predicts motion residual flows of object parts, and these flows are used to determine the articulation type and parameters. The network achieves an articulation type classification accuracy of 82.5% on novel object instances in trained categories. Experiments also show how this method enables generalization to novel categories and can be applied to real-world images without fine-tuning.