Multimodal integration learning of robot behavior using deep neural networks

Multimodal integration learning of robot behavior using deep neural networks
复制标题

DOI:
10.1016/j.robot.2014.03.003
复制
发表时间:
2014-06-01
影响因子:
4.3
通讯作者:
Ogata, Tetsuya
Ogata, Tetsuya
中科院分区:
计算机科学3区
文献类型:
--
作者:
Noda, Kuniaki;Arie, Hiroaki;Ogata, Tetsuya

文献摘要

被引文献

相似文献

对于人类准确地理解他们周围的世界,多模态集成是必不可少的,因为它提高了感知精度,减少了歧义。复制这种人类能力的计算模型可能有助于机器人在日常人类生活环境中的实际使用;然而,主要是因为传统机器学习算法遭受的可扩展性问题,机器人应用中的感觉运动信息处理通常通过模态依赖过程来实现。在本文中,我们提出了一种新的计算框架,能够整合感觉运动时间序列数据和基于深度学习方法的多模态融合表示的自组织。为了评估我们提出的模型,我们进行了两个行为学习实验,利用类人机器人,实验包括对象操作和响铃任务。从我们的实验结果,我们表明,大量的感觉运动信息,包括原始RGB图像,声音频谱和关节角度,直接融合生成更高层次的多模态表示。此外,我们证明了我们提出的框架实现了以下三个功能:(1)利用深度自动编码器的信息互补能力进行跨模态记忆检索;(2)利用多模态特征的泛化能力进行噪声鲁棒性行为识别;(3)基于获得的因果关系进行多模态因果关系获取和感觉运动预测。(C)2014作者由爱思唯尔公司出版
For humans to accurately understand the world around them, multimodal integration is essential because it enhances perceptual precision and reduces ambiguity. Computational models replicating such human ability may contribute to the practical use of robots in daily human living environments; however, primarily because of scalability problems that conventional machine learning algorithms suffer from, sensory-motor information processing in robotic applications has typically been achieved via modal-dependent processes. In this paper, we propose a novel computational framework enabling the integration of sensory-motor time-series data and the self-organization of multimodal fused representations based on a deep learning approach. To evaluate our proposed model, we conducted two behavior-learning experiments utilizing a humanoid robot; the experiments consisted of object manipulation and bell-ringing tasks. From our experimental results, we show that large amounts of sensory-motor information, including raw RGB images, sound spectrums, and joint angles, are directly fused to generate higher-level multimodal representations. Further, we demonstrated that our proposed framework realizes the following three functions: (1) cross-modal memory retrieval utilizing the information complementation capability of the deep autoencoder; (2) noise-robust behavior recognition utilizing the generalization capability of multimodal features; and (3) multimodal causality acquisition and sensory-motor prediction based on the acquired causality. (C) 2014 The Authors. Published by Elsevier B.V.