Human-robot interaction through real-time auditory and visual multiple-talker tracking

Human-robot interaction through real-time auditory and visual multiple-talker tracking
复制标题

通过实时听觉和视觉多说话者跟踪进行人机交互

DOI:
10.1109/iros.2001.977177
复制
发表时间:
2001
期刊:
Proceedings 2001 IEEE/RSJ International Conference on Intelligent Robots and Systems. Expanding the Societal Role of Robotics in the the Next Millennium (Cat. No.01CH37180)
影响因子:
--
通讯作者:
H. Kitano
H. Kitano
中科院分区:
--
文献类型:
--
作者:
HIroshi G. Okuno;K. Nakadai;K. Hidai;H. Mizoguchi;H. Kitano

文献摘要

被引文献

相似文献

Nakadai等人(2001)开发了一种实时听觉和视觉多说话者跟踪技术。在本文中,这种技术被应用到人机交互,包括接待员机器人和伴侣机器人在一个聚会。该系统包括人脸识别,语音识别,注意力集中控制,并在跟踪多个说话人的感觉运动任务。该系统是在一个上半身人形机器人上实现的,说话人跟踪是通过100 Base-TX网络连接的三个节点上的分布式处理来实现的。跟踪延迟为200 msec。注意力的焦点是通过使用声源方向和说话者位置作为线索来关联听觉和视觉流来控制的。一旦建立了关联,人形机器人就会将其面部保持在相关说话者的方向上。
Nakadai et al. (2001) have developed a real-time auditory and visual multiple-talker tracking technique. In this paper, this technique is applied to human-robot interaction including a receptionist robot and a companion robot at a party. The system includes face identification, speech recognition, focus-of-attention control, and sensorimotor task in tracking multiple talkers. The system is implemented on a upper-torso humanoid and the talker tracking is attained by distributed processing on three nodes connected by 100Base-TX network. The delay of tracking is 200 msec. Focus-of-attention is controlled by associating auditory and visual streams by using the sound source direction and talker position as a clue. Once an association is established, the humanoid keeps its face to the direction of the associated talker.