Learning Interpretable Control Dimensions for Speech Synthesis by Using External Data

Learning Interpretable Control Dimensions for Speech Synthesis by Using External Data
复制标题

DOI:
10.21437/interspeech.2018-2075
复制
发表时间:
2018-09
期刊:
ACM Transactions on Computer Systems (TOCS)
影响因子:
--
通讯作者:
Zack Hodari;O. Watts;S. Ronanki;Simon King
Zack Hodari;O. Watts;S. Ronanki;Simon King
中科院分区:
其他
文献类型:
--
作者:
Zack Hodari;O. Watts;S. Ronanki;Simon King

文献摘要

被引文献

相似文献

在创建文本到语音(TTS)系统时,我们可能希望控制语音的许多方面。我们提出了一个通用的方法,使控制的任意方面的讲话,我们证明了情感控制的任务。目前的TTS系统使用监督机器学习,因此严重依赖标记数据。如果没有标签可用于所需的控件维度,则创建可解释的控件变得具有挑战性。我们引入了一种方法,该方法使用外部标记数据(即不是用于训练声学模型的原始数据)来控制原始数据中未标记的维度。添加可解释的控制允许手动控制语音,以产生更吸引人的语音,用于有声读物等应用。我们使用听力测试来评估我们的方法。
There are many aspects of speech that we might want to control when creating text-to-speech (TTS) systems. We present a general method that enables control of arbitrary aspects of speech, which we demonstrate on the task of emotion control. Current TTS systems use supervised machine learning and are therefore heavily reliant on labelled data. If no labels are available for a desired control dimension, then creating interpretable control becomes challenging. We introduce a method that uses external, labelled data (i.e. not the original data used to train the acoustic model) to enable the control of dimensions that are not labelled in the original data. Adding interpretable control allows the voice to be manually controlled to produce more engaging speech, for applications such as audiobooks. We evaluate our method using a listening test.