Learning Paralinguistic Features from Audiobooks through Style Voice Conversion

Learning Paralinguistic Features from Audiobooks through Style Voice Conversion
复制标题

DOI:
10.18653/v1/2021.naacl-main.377
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Zakaria Aldeneh;Matthew Perez;Emily Mower Provost
Zakaria Aldeneh;Matthew Perez;Emily Mower Provost
中科院分区:
其他
文献类型:
--
作者:
Zakaria Aldeneh;Matthew Perez;Emily Mower Provost

文献摘要

相似文献

副语言是言语中的非词汇成分,在人与人的互动中起着至关重要的作用。设计用于识别非语言信息,特别是语音情感和风格的模型很难训练,因为可用的标记数据集有限。在这项工作中,我们提出了一个新的框架,使神经网络能够学习使用未注释情感的数据从语音中提取非语言属性。我们评估了学习的嵌入在情感识别和说话风格检测的下游任务上的效用,证明了表面声学特征以及从其他无监督方法提取的嵌入的显着改进。我们的工作使未来的系统能够利用学习嵌入提取器作为一个单独的组件,能够突出语音的非语言成分。
Paralinguistics, the non-lexical components of speech, play a crucial role in human-human interaction. Models designed to recognize paralinguistic information, particularly speech emotion and style, are difficult to train because of the limited labeled datasets available. In this work, we present a new framework that enables a neural network to learn to extract paralinguistic attributes from speech using data that are not annotated for emotion. We assess the utility of the learned embeddings on the downstream tasks of emotion recognition and speaking style detection, demonstrating significant improvements over surface acoustic features as well as over embeddings extracted from other unsupervised approaches. Our work enables future systems to leverage the learned embedding extractor as a separate component capable of highlighting the paralinguistic components of speech.