Speaker Adaptation of Various Components in Deep Neural Network based Speech Synthesis

Speaker Adaptation of Various Components in Deep Neural Network based Speech Synthesis
复制标题

DOI:
10.21437/ssw.2016-25
复制
发表时间:
2016-09
期刊:
--
影响因子:
--
通讯作者:
Shinji Takaki;Sangjin Kim;J. Yamagishi
Shinji Takaki;Sangjin Kim;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Shinji Takaki;Sangjin Kim;J. Yamagishi

文献摘要

相似文献

在本文中,我们研究了基于深度神经网络的语音合成中各种重要组件的说话人自适应的有效性,包括声学模型,声学特征提取和后滤波器。一般来说,说话者自适应技术,例如,用于Hacker的最大似然线性回归(MLLR)或用于DNN的学习隐藏单元贡献(LHUC)被应用于声学建模部分以改变语音特性或说话风格。然而,由于我们提出了一个基于多DNN的语音合成系统,其中多个组件基于前馈DNN来表示,因此说话人自适应技术不仅可以应用于声学建模部分,还可以应用于由DNN表示的其他组件。在使用少量自适应数据的实验中,我们基于LHUC进行自适应,并对基于DNN的声学模型、基于深度自动编码器的特征提取和基于DNN的后滤波器模型进行简单的额外微调,并将其与使用MLLR的基于HMM的语音合成系统进行比较。
In this paper, we investigate the effectiveness of speaker adaptation for various essential components in deep neural network based speech synthesis, including acoustic models, acoustic feature extraction, and post-filters. In general, a speaker adaptation technique, e.g., maximum likelihood linear regression (MLLR) for HMMs or learning hidden unit contributions (LHUC) for DNNs, is applied to an acoustic modeling part to change voice characteristics or speaking styles. However, since we have proposed a multiple DNN-based speech synthesis system, in which several components are represented based on feed-forward DNNs, a speaker adaptation technique can be applied not only to the acoustic modeling part but also to other components represented by DNNs. In experiments using a small amount of adaptation data, we performed adaptation based on LHUC and simple additional fine tuning for DNN-based acoustic models, deep auto-encoder based feature extraction, and DNN-based post-filter models and compared them with HMM-based speech synthesis systems using MLLR.