Cross-corpus acoustic emotion recognition from singing and speaking: A multi-task learning approach

Cross-corpus acoustic emotion recognition from singing and speaking: A multi-task learning approach
复制标题

DOI:
10.1109/icassp.2016.7472790
复制
发表时间:
2016-03
期刊:
2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Biqiao Zhang;E. Provost;Georg Essl
Biqiao Zhang;E. Provost;Georg Essl
中科院分区:
其他
文献类型:
--
作者:
Biqiao Zhang;E. Provost;Georg Essl

文献摘要

被引文献

相似文献

情感是通过语言和歌曲来表达的。先前的研究发现,虽然口头和歌唱的情绪识别是不同的任务,但它们是相关的。明确利用这种相关性的分类器可以实现比不这样做的分类器更好的性能。此外,语音情感识别的研究表明,当考虑性别时,情感的建模更准确。然而,目前还不清楚领域(语音或歌曲)和性别如何在情感识别系统中联合利用,也不清楚利用这些信息的系统如何在跨语料库设置中执行。在本文中,我们探索了一个多任务的情感识别框架,并比较了不同的分类模型和输出选择/融合方法,使用跨语料库评估的性能。我们的研究结果表明,当信息仅在密切相关的任务之间共享时,以及当不同模型的输出被融合时,分类准确率最高。
Emotion is expressed over both speech and song. Previous works have found that although spoken and sung emotion recognition are different tasks, they are related. Classifiers that explicitly utilize this relatedness can achieve better performance than classifiers that do not. Further, research in speech emotion recognition has demonstrated that emotion is more accurately modeled when gender is taken into account. However, it is not yet clear how domain (speech or song) and gender can be jointly leveraged in emotion recognition systems nor how systems leveraging this information can perform in cross-corpus settings. In this paper, we explore a multi-task emotion recognition framework and compare the performance across different classification models and output selection/fusion methods using cross-corpus evaluation. Our results show the classification accuracy is the highest when information is shared only between closely related tasks and when the output of disparate models are fused.