Robust speech interface based on audio and video information fusion for humanoid HRP-2

Robust speech interface based on audio and video information fusion for humanoid HRP-2
复制标题

基于音视频信息融合的人形HRP-2鲁棒语音接口

DOI:
10.1109/iros.2004.1389768
复制
发表时间:
2004
期刊:
2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE Cat. No.04CH37566)
影响因子:
--
通讯作者:
Kiyoshi Yamamoto
Kiyoshi Yamamoto
中科院分区:
--
文献类型:
--
作者:
Isao Hara;F. Asano;H. Asoh;J. Ogata;N. Ichimura;Y. Kawai;F. Kanehiro;H. Hirukawa;Kiyoshi Yamamoto

文献摘要

参考文献

被引文献

相似文献

在真实的世界中,人与机器人的协同工作需要机器人具有基于语音的交流功能。为了在嘈杂的真实的环境中实现这样的功能,机器人必须能够通过它们自己的资源从声音的混合物中提取人类所说的目标语音。提出了一种基于音视频信息融合的语音事件检测和提取方法。在该方法中,音频信息(使用麦克风阵列的声音定位)和视频信息(使用摄像机的人体跟踪)通过贝叶斯网络融合,以实现语音事件的检测。检测到的语音事件的信息,然后利用自适应波束形成的声音分离。在本文中,应用上述系统的人形机器人HRP-2的一些基本调查报告。输入设备,即麦克风阵列和摄像头,安装在HRP-2的头部,并对声音定位/分离性能的声学特性进行了研究。此外,人体跟踪系统进行了改进,使其可以在动态的情况下使用。最后通过离线实验对系统的整体性能进行了测试。
For cooperative work of robots and humans in the real world, a communicative function based on speech is indispensable for robots. To realize such a function in a noisy real environment, it is essential that robots be able to extract target speech spoken by humans from a mixture of sounds by their own resources. We have developed a method of detecting and extracting speech events based on the fusion of audio and video information. In this method, audio information (sound localization using a microphone array) and video information (human tracking using a camera) are fused by a Bayesian network to enable the detection of speech events. The information of detected speech events is then utilized in sound separation using adaptive beam forming. In this paper, some basic investigations for applying the above system to the humanoid robot HRP-2 are reported. Input devices, namely a microphone array and a camera, were mounted on the head of HRP-2, and acoustic characteristics for sound localization/separation performance were investigated. Also, the human tracking system was improved so that it can be used in a dynamic situation. Finally, overall performance of the system was tested via off-line experiments.
DOI: --
发表时间: 2003
期刊: Proceedings of The 6th International Conference on Information Fusion
影响因子: --
作者:
浅野 太;他
通讯作者: 他
利用音视频信息融合进行语音片段的检测和分离
DOI: --
发表时间: 2003
期刊: Proceedings of 8th European Conference on Speech Communication and Technology
影响因子: --
作者:
浅野 太;他
通讯作者: 他