Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech

Duration-Aware Pause Insertion Using Pre-Trained Language Model for Multi-Speaker Text-To-Speech
复制标题

DOI:
10.1109/icassp49357.2023.10096402
复制
发表时间:
2023-02
期刊:
ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
D. Yang;Tomoki Koriyama;Yuki Saito;Takaaki Saeki;Detai Xin;H. Saruwatari
D. Yang;Tomoki Koriyama;Yuki Saito;Takaaki Saeki;Detai Xin;H. Saruwatari
中科院分区:
其他
文献类型:
--
作者:
D. Yang;Tomoki Koriyama;Yuki Saito;Takaaki Saeki;Detai Xin;H. Saruwatari

文献摘要

相似文献

停顿插入,也被称为短语中断预测和短语处理,是TTS系统的重要组成部分,因为适当的自然持续时间的停顿可以显著提高合成语音的节奏和可理解性。然而,传统的短语模型忽略了不同说话人插入沉默停顿的不同风格,这可能会降低在多说话人语料库上训练的模型的性能。为此,我们提出了更强大的基于预训练语言模型的停顿插入框架。我们的方法使用在大规模文本语料库上预先训练的转换器双向编码器表示(BERT),注入说话人嵌入来捕获各种说话人特征。我们还利用时长感知暂停插入,以实现更自然的多扬声器TTS。我们开发并评估了两种类型的模型。第一种改进了传统的预测呼吸暂停(RPS)位置的语法模型,即在单词转换时没有标点符号的沉默停顿。它结合上下文信息进行说话人条件RP预测,并用来演示说话人信息对预测的影响。第二个模型进一步为基于音素的TTS模型设计,并执行持续时间感知的停顿插入,预测按持续时间分类的RP和标点符号指示的停顿(PIP)。评估结果表明,我们的模型提高了插入停顿的准确率和召回率,改善了合成语音的节奏。
Pause insertion, also known as phrase break prediction and phrasing, is an essential part of TTS systems because proper pauses with natural duration significantly enhance the rhythm and intelligibility of synthetic speech. However, conventional phrasing models ignore various speakers’ different styles of inserting silent pauses, which can degrade the performance of the model trained on a multi-speaker speech corpus. To this end, we propose more powerful pause insertion frameworks based on a pre-trained language model. Our approach uses bidirectional encoder representations from transformers (BERT) pre-trained on a large-scale text corpus, injecting speaker embeddings to capture various speaker characteristics. We also leverage duration-aware pause insertion for more natural multi-speaker TTS. We develop and evaluate two types of models. The first improves conventional phrasing models on the position prediction of respiratory pauses (RPs), i.e., silent pauses at word transitions without punctuation. It performs speaker-conditioned RP prediction considering contextual information and is used to demonstrate the effect of speaker information on the prediction. The second model is further designed for phoneme-based TTS models and performs duration-aware pause insertion, predicting both RPs and punctuation-indicated pauses (PIPs) that are categorized by duration. The evaluation results show that our models improve the precision and recall of pause insertion and the rhythm of synthetic speech.