Engagement recognition by a latent character model based on multimodal listener behaviors in spoken dialogue

Engagement recognition by a latent character model based on multimodal listener behaviors in spoken dialogue
复制标题

基于口语对话中多模态听众行为的潜在角色模型的参与识别

DOI:
10.1017/atsip.2018.11
复制
发表时间:
2018
期刊:
APSIPA Trans. Signal & Information Processing
影响因子:
--
通讯作者:
and T.Kawahara
and T.Kawahara
中科院分区:
--
文献类型:
--
作者:
K.Inoue;D.Lala;K.Takanashi;and T.Kawahara

文献摘要

相似文献

用户粘性代表用户对当前对话感兴趣并愿意继续对话的程度。参与识别将为对话系统生成用户自适应行为提供重要线索。本文研究了基于反向通道、大笑、点头和注视等多模态听者行为的参与识别。在敬业度的注释中,由于敬业度感知的主观性,每个注释者的基础事实数据往往不同。为了解决这个问题,我们假设每个注释者都有一个潜在的性格,会影响他/她对参与的感知。我们提出了一个分层贝叶斯模型,该模型估计每个注释者的参与和特征作为潜在变量。此外,我们将参与识别模型与听者行为自动检测相结合,实现在线参与识别。实验结果表明,与不考虑多数投票等特征的其他方法相比,该模型提高了识别精度。我们还在不降低准确性的情况下实现了在线参与识别。
Engagement represents how much a user is interested in and willing to continue the current dialogue. Engagement recognition will provide an important clue for dialogue systems to generate adaptive behaviors for the user. This paper addresses engagement recognition based on multimodal listener behaviors of backchannels, laughing, head nodding, and eye gaze. In the annotation of engagement, the ground-truth data often differs from one annotator to another due to the subjectivity of the perception of engagement. To deal with this, we assume that each annotator has a latent character that affects his/her perception of engagement. We propose a hierarchical Bayesian model that estimates both engagement and the character of each annotator as latent variables. Furthermore, we integrate the engagement recognition model with automatic detection of the listener behaviors to realize online engagement recognition. Experimental results show that the proposed model improves recognition accuracy compared with other methods which do not consider the character such as majority voting. We also achieve online engagement recognition without degrading accuracy.