On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis

On the Interplay between Sparsity, Naturalness, Intelligibility, and Prosody in Speech Synthesis
复制标题

DOI:
10.1109/icassp43922.2022.9747728
复制
发表时间:
2021-10
期刊:
ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Cheng-I Lai;Erica Cooper;Yang Zhang;Shiyu Chang;Kaizhi Qian;Yiyuan Liao;Yung-Sung Chuang;Alexander H. Liu;J. Yamagishi;David Cox;James R. Glass
Cheng-I Lai;Erica Cooper;Yang Zhang;Shiyu Chang;Kaizhi Qian;Yiyuan Liao;Yung-Sung Chuang;Alexander H. Liu;J. Yamagishi;David Cox;James R. Glass
中科院分区:
其他
文献类型:
--
作者:
Cheng-I Lai;Erica Cooper;Yang Zhang;Shiyu Chang;Kaizhi Qian;Yiyuan Liao;Yung-Sung Chuang;Alexander H. Liu;J. Yamagishi;David Cox;James R. Glass

文献摘要

相似文献

端到端的文本到语音(TTS)模型是否过度参数化?这些模型可以被修剪到什么程度,它们的合成能力会发生什么变化?这项工作作为一个起点,探索修剪频谱预测网络和声码器。我们彻底调查稀疏性及其后续影响之间的权衡合成语音。此外,我们还探讨了TTS修剪的几个方面:微调数据量与稀疏性,TTS增强利用未说出的文本,并结合知识蒸馏和修剪。我们的研究结果表明,不仅是端到端的TTS模型高度prunable,而且,也许令人惊讶的是,修剪TTS模型可以产生合成语音具有相同或更高的自然度和可理解性,具有类似的韵律。我们所有的实验都是在公开的模型上进行的,这项工作的发现得到了大规模主观测试和客观测量的支持。代码和200修剪模型,以方便未来的研究在TTS1的效率。
Are end-to-end text-to-speech (TTS) models over-parametrized? To what extent can these models be pruned, and what happens to their synthesis capabilities? This work serves as a starting point to explore pruning both spectrogram prediction networks and vocoders. We thoroughly investigate the tradeoffs between sparsity and its subsequent effects on synthetic speech. Additionally, we explore several aspects of TTS pruning: amount of finetuning data versus sparsity, TTS-Augmentation to utilize unspoken text, and combining knowledge distillation and pruning. Our findings suggest that not only are end-to-end TTS models highly prunable, but also, perhaps surprisingly, pruned TTS models can produce synthetic speech with equal or higher naturalness and intelligibility, with similar prosody. All of our experiments are conducted on publicly available models, and findings in this work are backed by large-scale subjective tests and objective measures. Code and 200 pruned models are made available to facilitate future research on efficiency in TTS1.