Deep Gated Multi-modal Learning: In-hand Object Pose Changes Estimation using Tactile and Image Data

Deep Gated Multi-modal Learning: In-hand Object Pose Changes Estimation using Tactile and Image Data
复制标题

深度门控多模态学习:使用触觉和图像数据估计手中物体姿势变化

DOI:
--
复制
发表时间:
2019
期刊:
IEEE/RJS International Conference on Intelligent RObots and Systems
影响因子:
--
通讯作者:
K. Takahashi
K. Takahashi
中科院分区:
--
文献类型:
--
作者:
T. Anzai;K. Takahashi

文献摘要

被引文献

相似文献

对于手内操作来说,手内物体位姿的估计是将物体操作到目标位姿的重要功能之一。由于手内操作往往会导致手或物体本身的遮挡,因此仅图像信息不足以进行手内物体姿态估计。在这种情况下可以使用多种模态,优点是其他模态可以补偿阻塞、噪声和传感器故障。即使确定与情况相对应的模态的利用率(称为可靠性值)是重要的,但是这种模型的手动设计是困难的,特别是对于各种情况。在本文中,我们提出了深度门控多模态学习,它通过端到端的深度学习来自我确定每个模态的可靠性值。在实验中,RGB摄像机和GelSight触觉传感器连接到索耶机器人的平行夹持器,并估计在抓取过程中的对象位姿变化。在实验中总共使用了15个物体。在该模型中,模态的可靠性值被确定根据噪声水平和故障的每一个模态,它被证实,即使是未知对象的姿态变化估计。1
For in-hand manipulation, estimation of the object pose inside the hand is one of the important functions to manipulate objects to the target pose. Since in-hand manipulation tends to cause occlusions by the hand or the object itself, image information only is not sufficient for in-hand object pose estimation. Multiple modalities can be used in this case, the advantage is that other modalities can compensate for occlusion, noise, and sensor malfunctions. Even though deciding the utilization rate of a modality (referred to as reliability value) corresponding to the situations is important, the manual design of such models is difficult, especially for various situations. In this paper, we propose deep gated multi-modal learning, which self-determines the reliability value of each modality through end-to-end deep learning. For the experiments, an RGB camera and a GelSight tactile sensor were attached to the parallel gripper of the Sawyer robot, and the object pose changes were estimated during grasping. A total of 15 objects were used in the experiments. In the proposed model, the reliability values of the modalities were determined according to the noise level and failure of each modality, and it was confirmed that the pose change was estimated even for unknown objects. 1