Learning Interpretable Control Dimensions for Speech Synthesis by Using External Data
Learning Interpretable Control Dimensions for Speech Synthesis by Using External Data
复制标题
DOI:
10.21437/interspeech.2018-2075
复制
发表时间:
2018-09
期刊:
影响因子:
--
通讯作者:
Zack Hodari;O. Watts;S. Ronanki;Simon King
中科院分区:
文献类型:
--
作者:
Zack Hodari;O. Watts;S. Ronanki;Simon King
There are many aspects of speech that we might want to control when creating text-to-speech (TTS) systems. We present a general method that enables control of arbitrary aspects of speech, which we demonstrate on the task of emotion control. Current TTS systems use supervised machine learning and are therefore heavily reliant on labelled data. If no labels are available for a desired control dimension, then creating interpretable control becomes challenging. We introduce a method that uses external, labelled data (i.e. not the original data used to train the acoustic model) to enable the control of dimensions that are not labelled in the original data. Adding interpretable control allows the voice to be manually controlled to produce more engaging speech, for applications such as audiobooks. We evaluate our method using a listening test.