Domain Adversarial for Acoustic Emotion Recognition

Domain Adversarial for Acoustic Emotion Recognition
复制标题

DOI:
10.1109/taslp.2018.2867099
复制
发表时间:
2018-12-01
影响因子:
5.4
通讯作者:
Busso, Carlos
Busso, Carlos
中科院分区:
计算机科学2区
文献类型:
--
作者:
Abdelwahab, Mohammed;Busso, Carlos

文献摘要

被引文献

相似文献

语音情感识别的性能受到用于构建和评估模型的训练集(源域)和测试集(目标域)之间数据分布差异的影响。这是一个常见问题,因为多项研究表明,当情感分类器接触到与用于构建情感分类器的分布不匹配的数据时,情感分类器的性能会下降。当训练和测试数据来自不同领域时,数据分布的差异变得非常明显,导致开发和测试性能之间存在较大的性能差距。由于注释新数据的成本高昂且未标记数据丰富,因此从可用的未标记数据中提取尽可能多的有用信息至关重要。这项研究探讨了使用对抗性多任务训练来提取训练域和测试域之间的共同表示。主要任务是预测基于情感属性的唤醒、效价或支配性描述符。第二个任务是学习一个通用的表示,其中训练域和测试域无法区分。通过使用梯度反转层,来自域分类器的梯度用于使源域表示和目标域表示更接近。我们表明,利用未标记的数据始终可以在所有情感维度上带来更好的情感识别性能。我们可视化对抗性训练对所提出的深度学习架构的特征表示的影响。分析表明,随着数据传递到网络的更深层,训练域和测试域的数据表示会收敛。我们还评估了使用浅层神经网络与深层神经网络时的性能差异,以及任务和域分类器使用的共享层数量的影响。
The performance of speech emotion recognition is affected by the differences in data distributions between train (source domain) and test (target domain) sets used to build and evaluate the models. This is a common problem, as multiple studies have shown that the performance of emotional classifiers drops when they are exposed to data that do not match the distribution used to build the emotion classifiers. The difference in data distributions becomes very clear when the training and testing data come from different domains, causing a large performance gap between development and testing performance. Due to the high cost of annotating new data and the abundance of unlabeled data, it is crucial to extract as much useful information as possible from the available unlabeled data. This study looks into the use of adversarial multitask training to extract a common representation between train and test domains. The primary task is to predict emotionalattribute-based descriptors for arousal, valence, or dominance. The secondary task is to learn a common representation, where the train and test domains cannot be distinguished. By using a gradient reversal layer, the gradients coming from the domain classifier are used to bring the source and target domain representations closer. We show that exploiting unlabeled data consistently leads to better emotion recognition performance across all emotional dimensions. We visualize the effect of adversarial training on the feature representation across the proposed deep learning architecture. The analysis shows that the data representations for the train and test domains converge as the data are passed to deeper layers of the network. We also evaluate the difference in performance when we use a shallow neural network versus a deep neural network and the effect of the number of shared layers used by the task and domain classifiers.