Multimodal addressee detection in multiparty dialogue systems

Multimodal addressee detection in multiparty dialogue systems
复制标题

多方对话系统中的多模式收件人检测

DOI:
--
复制
发表时间:
2015
期刊:
IEEE International Conference on Acoustics, Speech, and Signal Processing
影响因子:
--
通讯作者:
M. Slaney
M. Slaney
中科院分区:
--
文献类型:
--
作者:
T. Tsai;A. Stolcke;M. Slaney

文献摘要

被引文献

相似文献

收件人检测回答问题,“你在和我说话吗?”当多个用户与对话系统交互时,重要的是要知道用户何时对计算机说话以及他或她何时对另一个人说话。我们从多模态的角度来处理这个问题,使用词汇,声学,视觉,对话状态和波束形成信息。使用多方对话系统的数据,我们证明了使用多种方式的好处,使用单一的方式。我们还评估的相对重要性的各种方式在预测收件人。在我们的实验中,我们发现声学特征是迄今为止最重要的,ASR和系统状态信息是有用的,视觉和波束形成功能提供了一些额外的好处。我们的研究表明,声学,词汇,和系统状态信息是一个有效的,经济的组合方式使用的收件人检测。
Addressee detection answers the question, “Are you talking to me?” When multiple users interact with a dialogue system, it is important to know when a user is speaking to the computer and when he or she is speaking to another person. We approach this problem from a multimodal perspective, using lexical, acoustic, visual, dialog state, and beam-forming information. Using data from a multiparty dialogue system, we demonstrate the benefit of using multiple modalities over using a single modality. We also assess the relative importance of the various modalities in predicting the addressee. In our experiments, we find that acoustic features are by far the most important, that ASR and system-state information are useful, and that visual and beamforming features provide little additional benefit. Our study suggests that acoustic, lexical, and system state information are an effective, economical combination of modalities to use in addressee detection.