Deep Multi-Sensory Object Category Recognition Using Interactive Behavioral Exploration

Deep Multi-Sensory Object Category Recognition Using Interactive Behavioral Exploration
复制标题

DOI:
10.1109/icra.2019.8794095
复制
发表时间:
2019-05
期刊:
2019 International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Gyan Tatiya;Jivko Sinapov
Gyan Tatiya;Jivko Sinapov
中科院分区:
其他
文献类型:
--
作者:
Gyan Tatiya;Jivko Sinapov

文献摘要

被引文献

相似文献

在识别物体及其属性时,人类会使用操纵该物体时产生的多种感官模式的特征。受这种认知过程的启发,我们提出了一种用于物体类别识别的深度学习方法,该方法使用视觉、听觉和触觉感官数据以及探索行为(例如抓握、举起、推动等)。在我们的方法中,当机器人对物体执行动作时,它使用张量训练门控循环单元网络来处理其视觉数据,并使用卷积神经网络来处理触觉和听觉数据。我们提出了一种新颖的策略来训练输入视频、音频和触觉数据的单个神经网络,并证明其性能优于每种感觉模态的单独神经网络。该方法在一个数据集上进行了评估,其中机器人探索了 100 个不同的物体,每个物体属于 20 个类别之一。虽然视觉信息是大多数类别的主要形式,但添加额外的触觉和听觉网络进一步提高了机器人的类别识别准确性。对于某些行为,我们的方法优于之前发布的数据集基线,该基线对每种模态使用手工制作的特征。我们还表明,机器人不需要整个交互过程中的感官数据,而是可以在行为执行的早期做出良好的预测。
When identifying an object and its properties, humans use features from multiple sensory modalities produced when manipulating the object. Motivated by this cognitive process, we propose a deep learning methodology for object category recognition which uses visual, auditory, and haptic sensory data coupled with exploratory behaviors (e.g., grasping, lifting, pushing, etc.). In our method, as the robot performs an action on an object, it uses a Tensor-Train Gated Recurrent Unit network to process its visual data, and Convolutional Neural Networks to process haptic and auditory data. We propose a novel strategy to train a single neural network that inputs video, audio and haptic data, and demonstrate that its performance is better than separate neural networks for each sensory modality. The proposed method was evaluated on a dataset in which the robot explored 100 different objects, each belonging to one of 20 categories. While the visual information was the dominant modality for most categories, adding the additional haptic and auditory networks further improves the robot’s category recognition accuracy. For some of the behaviors, our approach outperforms the previous published baseline for the dataset which used handcrafted features for each modality. We also show that a robot does not need the sensory data from the entire interaction, but instead can make a good prediction early on during behavior execution.