Cross-Corpus Acoustic Emotion Recognition with Multi-Task Learning: Seeking Common Ground While Preserving Differences

Cross-Corpus Acoustic Emotion Recognition with Multi-Task Learning: Seeking Common Ground While Preserving Differences
复制标题

DOI:
10.1109/taffc.2017.2684799
复制
发表时间:
2019
影响因子:
11.2
通讯作者:
Biqiao Zhang;E. Provost;Georg Essl
Biqiao Zhang;E. Provost;Georg Essl
中科院分区:
计算机科学2区
文献类型:
--
作者:
Biqiao Zhang;E. Provost;Georg Essl

文献摘要

被引文献

相似文献

由于情感识别在许多应用中的潜力,人们对它的兴趣越来越大。然而,一个普遍的挑战是由诸如语料库之间的差异、说话者的性别和表达的“域”(例如,无论表达是说出来的还是唱出来的)。以前的工作已经解决了这一挑战,结合跨语料库和/或性别的数据,或明确控制这些因素。在这项工作中,我们调查语料库,领域,和性别的跨语料库的情感识别系统的泛化能力的影响。我们使用多任务学习方法,根据这些因素定义任务。我们发现,通过多任务学习将语料库,域和性别引起的变化优于将任务视为相同或独立的方法。对于多域数据,域是比性别更大的区分因素。当只考虑语音域时,性别和语料库同样有影响力。用性别来定义任务比用语料库或语料库和性别来定义配价更有利,而用激活来定义则相反。平均而言,跨语料库性能随着训练语料库的数量而增加。结果表明,有效的跨语料库建模需要我们了解情感表达模式如何作为非情感因素的函数而变化。
There is growing interest in emotion recognition due to its potential in many applications. However, a pervasive challenge is the presence of data variability caused by factors such as differences across corpora, speaker’s gender, and the “domain” of expression (e.g., whether the expression is spoken or sung). Prior work has addressed this challenge by combining data across corpora and/or genders, or by explicitly controlling for these factors. In this work, we investigate the influence of corpus, domain, and gender on the cross-corpus generalizability of emotion recognition systems. We use a multi-task learning approach, where we define the tasks according to these factors. We find that incorporating variability caused by corpus, domain, and gender through multi-task learning outperforms approaches that treat the tasks as either identical or independent. Domain is a larger differentiating factor than gender for multi-domain data. When considering only the speech domain, gender and corpus are similarly influential. Defining tasks by gender is more beneficial than by either corpus or corpus and gender for valence, while the opposite holds for activation. On average, cross-corpus performance increases with the number of training corpora. The results demonstrate that effective cross-corpus modeling requires that we understand how emotion expression patterns change as a function of non-emotional factors.