Semi-Supervised Speech Emotion Recognition With Ladder Networks

Semi-Supervised Speech Emotion Recognition With Ladder Networks
复制标题

DOI:
10.1109/taslp.2020.3023632
复制
发表时间:
2020-01-01
影响因子:
5.4
通讯作者:
Busso, Carlos
Busso, Carlos
中科院分区:
计算机科学2区
文献类型:
--
作者:
Parthasarathy, Srinivas;Busso, Carlos

文献摘要

被引文献

相似文献

语音情感识别(SER)系统在诸如医疗保健、教育以及安全和国防的各种领域中找到应用。这些系统的一个主要缺点是它们在不同条件下缺乏泛化能力。例如,在某些数据库上表现出上级性能的系统在其他语料库上测试时表现出较差的性能。这个问题可以通过在来自目标域的大量标记数据上训练模型来解决,这是昂贵且耗时的。另一种方法是增加模型的通用性。实现这一目标的一种有效方法是通过多任务学习(MTL)来正则化模型,其中辅助任务与主要任务一起沿着学习。这些方法通常需要使用标记的数据,这是计算昂贵的收集情感识别(性别,说话人身份,年龄或其他情感描述符)。这项研究提出了使用梯形网络的情感识别,它利用了无监督的辅助任务。主要任务是预测情感属性的回归问题。辅助任务是使用去噪自动编码器重建中间特征表示。这个辅助任务不需要标签,因此可以用来自目标域的大量未标记数据以半监督的方式训练框架。这项研究表明,所提出的方法创建了一个强大的框架SER,实现上级性能比完全监督单任务学习(STL)和MTL基线。我们实现的方法与文件级或帧级的功能,展示了我们的方法的灵活性。此外,在跨语料库的设置中,使用层次特征评估梯形网络的泛化,获得重要的改进。与STL基线相比,所提出的方法在语料库内评估的一致性相关系数(CCC)方面实现了3.0%和3.5%之间的相对增益,在跨语料库评估方面实现了16.1%和74.1%之间的相对增益,突出了架构的强大功能。
Speech emotion recognition (SER) systems find applications in various fields such as healthcare, education, and security and defense. A major drawback of these systems is their lack of generalization across different conditions. For example, systems that show superior performance on certain databases show poor performance when tested on other corpora. This problem can be solved by training models on large amounts of labeled data from the target domain, which is expensive and time-consuming. Another approach is to increase the generalization of the models. An effective way to achieve this goal is by regularizing the models through multitask learning (MTL), where auxiliary tasks are learned along with the primary task. These methods often require the use of labeled data which is computationally expensive to collect for emotion recognition (gender, speaker identity, age or other emotional descriptors). This study proposes the use of ladder networks for emotion recognition, which utilizes an unsupervised auxiliary task. The primary task is a regression problem to predict emotional attributes. The auxiliary task is the reconstruction of intermediate feature representations using a denoising autoencoder. This auxiliary task does not require labels so it is possible to train the framework in a semi-supervised fashion with abundant unlabeled data from the target domain. This study shows that the proposed approach creates a powerful framework for SER, achieving superior performance than fully supervised single-task learning (STL) and MTL baselines. We implement the approach with sentence-level or frame-level features, demonstrating the flexibility of our approach. Additionally, the generalization of the ladder networks is evaluated in cross-corpus settings using sentence-level features, obtaining important improvements. Compared to the STL baselines, the proposed approach achieves relative gains in concordance correlation coefficient (CCC) between 3.0% and 3.5% for within corpus evaluations, and between 16.1% and 74.1% for cross corpus evaluations, highlighting the power of the architecture.