Retrieving Speech Samples with Similar Emotional Content Using a Triplet Loss Function

Retrieving Speech Samples with Similar Emotional Content Using a Triplet Loss Function
复制标题

DOI:
10.1109/icassp.2019.8683273
复制
发表时间:
2019-05
期刊:
ICASSP 2019 - 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
John Harvill;Mohammed Abdel-Wahab;Reza Lotfian;C. Busso
John Harvill;Mohammed Abdel-Wahab;Reza Lotfian;C. Busso
中科院分区:
其他
文献类型:
--
作者:
John Harvill;Mohammed Abdel-Wahab;Reza Lotfian;C. Busso

文献摘要

被引文献

相似文献

识别具有相似情感内容的语音的能力对于许多应用是有价值的,包括语音检索、监视和情感语音合成。虽然目前基于分类或回归的语音情感识别公式不适合这项任务,但基于偏好学习的解决方案为这项任务提供了有吸引力的方法。本文的目的是找到语音样本,情感相似的锚语音样本作为查询提供。这种新的配方打开了有趣的研究问题。一台机器能多好地完成这项任务?自动算法的准确性与人类执行此任务的性能相比如何?这项研究通过使用三元组损失函数训练深度学习模型来解决这些问题,将声学特征映射到对该任务具有区分性的嵌入中。网络接收锚语音样本和两个竞争语音样本,并且任务是确定候选语音样本中的哪一个传达与由锚传达的情感最接近的情感内容。通过比较我们的模型的结果与人类的感知评价,这项研究表明,所提出的方法在检索具有相似情感内容的样本时具有非常接近人类的性能。
The ability to identify speech with similar emotional content is valuable to many applications, including speech retrieval, surveillance, and emotional speech synthesis. While current formulations in speech emotion recognition based on classification or regression are not appropriate for this task, solutions based on preference learning offer appealing approaches for this task. This paper aims to find speech samples that are emotionally similar to an anchor speech sample provided as a query. This novel formulation opens interesting research questions. How well can a machine complete this task? How does the accuracy of automatic algorithms compare to the performance of a human performing this task? This study addresses these questions by training a deep learning model using a triplet loss function, mapping the acoustic features into an embedding that is discriminative for this task. The network receives an anchor speech sample and two competing speech samples, and the task is to determine which of the candidate speech sample conveys the closest emotional content to the emotion conveyed by the anchor. By comparing the results from our model with human perceptual evaluations, this study demonstrates that the proposed approach has performance very close to human performance in retrieving samples with similar emotional content.