Prosody Prediction from Syntactic, Lexical, and Word Embedding Features

Prosody Prediction from Syntactic, Lexical, and Word Embedding Features
复制标题

DOI:
10.21437/ssw.2019-48
复制
发表时间:
2019-09
期刊:
10th ISCA Workshop on Speech Synthesis (SSW 10)
影响因子:
--
通讯作者:
Rose Sloan;S. S. Akhtar-S.;Bryan Li;Ritvik Shrivastava;Agustin Gravano;Julia Hirschberg
Rose Sloan;S. S. Akhtar-S.;Bryan Li;Ritvik Shrivastava;Agustin Gravano;Julia Hirschberg
中科院分区:
其他
文献类型:
--
作者:
Rose Sloan;S. S. Akhtar-S.;Bryan Li;Ritvik Shrivastava;Agustin Gravano;Julia Hirschberg

文献摘要

被引文献

相似文献

从文本中准确地预测韵律会使TTS听起来更自然。在这项工作中,我们使用了一组新的特征来预测文本中的托比基音和短语边界。我们研究了各种各样的基于文本的特征,包括许多新的句法特征,几种类型的词嵌入,共指特征,LIWC特征,以及特定fi城市信息。我们的工作重点放在波士顿广播新闻语料库上,这是一个TOBI标记的相对干净的新闻广播语料库,但也在AuDIX(一个较小的已读新闻语料库)和哥伦比亚游戏语料库(一个会话语音语料库)上测试我们的Classifier,以便测试我们的模型在跨语料库设置中的适用性。我们的结果显示了在这两个任务上的强大性能,以及我们的模型在跨语料库应用方面的一些有希望的结果。
Accurate prosody prediction from text leads to more natural-sounding TTS. In this work, we employ a new set of features to predict ToBI pitch accent and phrase boundaries from text. We investigate a wide variety of text-based features, including many new syntactic features, several types of word embeddings, co-reference features, LIWC features, and specificity information. We focus our work on the Boston Radio News Corpus, a ToBI-labeled corpus of relatively clean news broad-casts, but also test our classifiers on Audix, a smaller corpus of read news, and on the Columbia Games Corpus, a corpus of conversational speech, in order to test the applicability of our model in cross-corpus settings. Our results show strong performance on both tasks, as well as some promising results for cross-corpus applications of our models.