Recognizing Social Signals with Weakly Supervised Multitask Learning for Multimodal Dialogue Systems

Recognizing Social Signals with Weakly Supervised Multitask Learning for Multimodal Dialogue Systems
复制标题

DOI:
10.1145/3462244.3479927
复制
发表时间:
2021-10
期刊:
Proceedings of the 2021 International Conference on Multimodal Interaction
影响因子:
--
通讯作者:
Yuki Hirano;S. Okada;Kazunori Komatani
Yuki Hirano;S. Okada;Kazunori Komatani
中科院分区:
其他
文献类型:
--
作者:
Yuki Hirano;S. Okada;Kazunori Komatani

文献摘要

相似文献

社会信号处理是一种用于从言语和非言语多模态信息中推断人类内在状态(包括态度、情感和印象)的方法。训练社会信号识别模型的困难在于,由于对情感等社会信号的标注是一项主观且模糊的任务,多个编码员给出的真实(目标)标签往往不一致。我们将弱监督学习(WSL)算法引入到这种目标标签不一定准确的不准确监督环境中。本文的新挑战是探索一种有效的用于识别社会信号的WSL策略。该策略通过在人机对话环境中收集的包括音频、视觉和语言数据的两个多模态数据集进行了验证。首先,我们阐明所提出的用于深度神经网络(DNN)的WSL策略(称为三元教学)在几乎所有分类任务中都效果良好。其次,我们证明了整合WSL和多任务学习(MTL)的有效性,多任务学习利用了数据集中的几种标签类型。第三,我们表明,在跨语料库环境中,我们提出的方法比现有的用于DNN的训练算法(课程学习)的精度下降更少,最大改进幅度为7.2%。
Social signal processing is a methodology that is used to infer human inner states, including attitudes, sentiments and impressions, from verbal and nonverbal multimodal information. The difficulty in training a social signal recognition model is that the ground-truth (target) labels given by multiple coders often disagree because the annotation of social signals such as sentiments is a subjective and ambiguous task. We introduce weakly supervised learning (WSL) algorithms to such an inaccurate supervision setting in which the target label is not necessarily accurate. The novel challenge in this paper is to explore an effective WSL strategy for recognizing social signals. The strategy is verified through two multimodal datasets including audio, visual, and linguistic data collected in a human-agent dialogue setting. First, we clarify that the proposed WSL strategy for deep neural networks (DNNs), called tri-teaching works well in almost all classification tasks. Second, we demonstrate the effectiveness of integrating WSL and multitask learning (MTL), which exploits several label types in the datasets. Third, we show that our proposed approach achieves less accuracy degradation than an existing training algorithm for a DNN (curriculum learning) in a cross-corpus setting, with a maximum improvement of 7.2%.