Human–robot non-verbal interaction empowered by real-time auditory and visual multiple-talker tracking

Human–robot non-verbal interaction empowered by real-time auditory and visual multiple-talker tracking
复制标题

实时听觉和视觉多说话者跟踪支持人机非语言交互

DOI:
10.1163/156855303321165088
复制
发表时间:
2003
期刊:
影响因子:
2
通讯作者:
H. Kitano
H. Kitano
中科院分区:
计算机科学4区
文献类型:
--
作者:
HIroshi G. Okuno;K. Nakadai;K. Hidai;H. Mizoguchi;H. Kitano

文献摘要

被引文献

相似文献

声音对于增强视觉体验和人机交互至关重要,但大多数研发工作通常主要针对声音生成、语音合成和语音识别。听觉场景分析之所以很少受到关注,是因为实时感知混合声音很困难。最近,Nakadai 等人。开发了实时听觉和视觉多说话者跟踪技术。在本文中,该技术应用于人与机器人的言语和非言语交互,包括派对上的接待机器人和陪伴机器人。该系统包括面部识别、语音识别、注意力焦点控制和跟踪多个说话者的感觉运动任务。该系统是在一个名为 SIG 的上躯干人形机器人上实现的,并且通过 100Base-TX 网络连接的三个节点上的分布式处理来实现说话者跟踪。跟踪总体延迟为200ms。通过使用声源方向和说话者位置作为线索,将听觉和视觉流关联起来,来控制注意力焦点。一旦建立关联,类人生物就会将脸朝向关联说话者的方向。
Sound is essential to enhance visual experience and human-robot interaction, but most research and development efforts are usually made mainly towards sound generation, speech synthesis and speech recognition. The reason why only little attention has been paid to auditory scene analysis is that real-time perception of a mixture of sounds is difficult. Recently, Nakadai et al. have developed real-time auditory and visual multiple-talker tracking technology. In this paper, this technology is applied to human—robot verbal and non-verbal interaction including a receptionist robot and a companion robot at a party. The system includes face identification, speech recognition, focus-of-attention control and a sensorimotor task in tracking multiple talkers. The system is implemented on an upper-torso humanoid called SIG and the talker tracking is attained by distributed processing on three nodes connected by a 100Base-TX network. The overall delay of tracking is 200 ms. Focus-of-attention is controlled by associating auditory and visual streams with using the sound source direction and talker position as a clue. Once an association is established, the humanoid keeps its face towards the direction of the associated talker.