Shared acoustic codes underlie emotional communication in music and speech-Evidence from deep transfer learning.

Shared acoustic codes underlie emotional communication in music and speech-Evidence from deep transfer learning.
复制标题

DOI:
10.1371/journal.pone.0179289
复制
发表时间:
2017
期刊:
影响因子:
3.7
通讯作者:
Schuller B
Schuller B
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Coutinho E;Schuller B

文献摘要

被引文献

相似文献

音乐和语言在声学领域的情感交流方面表现出惊人的相似之处,在这种方式下,特定情感的交流至少在一定程度上是通过共享的声学模式来实现的。从情感科学的观点来看,确定这两个领域之间的重叠程度是理解这种现象背后的共同机制的基础。从机器学习的角度来看,音乐和语音中表达情感的声学代码之间的重叠为扩大可用于开发音乐和语音情感识别系统的数据量打开了新的可能性。在这篇文章中,我们研究了音乐和言语中情绪(唤醒和配价)的时间连续预测,以及这些领域之间的迁移学习。我们建立了一个比较框架,包括内部实验(即在同一通道上训练和测试的模型,无论是音乐还是语音)和跨域实验(即在一种通道上训练的模型和在另一通道上测试的模型)。在跨域的背景下,我们评估了两种策略-域间直接转移和转移学习技术(基于去噪自动编码器的特征表示转移)对缩小特征空间分布差距的贡献。我们的结果表明,在有和没有两个方向的特征表示转移的情况下,具有良好的跨域泛化性能。在音乐的情况下,跨域方法对价估计的性能优于域内模型,而对于语音,域内模型的性能最好。这是第一次在时间连续的域中展示用于音乐和语音中情感表达的共享声学代码。
Music and speech exhibit striking similarities in the communication of emotions in the acoustic domain, in such a way that the communication of specific emotions is achieved, at least to a certain extent, by means of shared acoustic patterns. From an Affective Sciences points of view, determining the degree of overlap between both domains is fundamental to understand the shared mechanisms underlying such phenomenon. From a Machine learning perspective, the overlap between acoustic codes for emotional expression in music and speech opens new possibilities to enlarge the amount of data available to develop music and speech emotion recognition systems. In this article, we investigate time-continuous predictions of emotion (Arousal and Valence) in music and speech, and the Transfer Learning between these domains. We establish a comparative framework including intra- (i.e., models trained and tested on the same modality, either music or speech) and cross-domain experiments (i.e., models trained in one modality and tested on the other). In the cross-domain context, we evaluated two strategies—the direct transfer between domains, and the contribution of Transfer Learning techniques (feature-representation-transfer based on Denoising Auto Encoders) for reducing the gap in the feature space distributions. Our results demonstrate an excellent cross-domain generalisation performance with and without feature representation transfer in both directions. In the case of music, cross-domain approaches outperformed intra-domain models for Valence estimation, whereas for Speech intra-domain models achieve the best performance. This is the first demonstration of shared acoustic codes for emotional expression in music and speech in the time-continuous domain.