Corpus-based generation of prosodic features from text based on generation process model

Corpus-based generation of prosodic features from text based on generation process model
复制标题

DOI:
10.21437/interspeech.2007-228
复制
发表时间:
2007
期刊:
影响因子:
6.1
通讯作者:
K. Hirose;K. Ochi;N. Minematsu
K. Hirose;K. Ochi;N. Minematsu
中科院分区:
医学2区
文献类型:
--
作者:
K. Hirose;K. Ochi;N. Minematsu

文献摘要

相似文献

构建了一个从文本输入中提取韵律特征的总体方案。该方法包括基于语料库的预测停顿,电话持续时间和基本频率(F0),在这个顺序,和信息预测在较早的过程中被利用在下面的过程。由于F0的预测是根据F0轮廓生成过程模型的命令值而不是直接的F0值进行的,所以可以稳定和灵活地控制F0轮廓。通过在重音命令定时上添加约束作为后处理,当使用由该方法生成的韵律特征合成语音时,实现了更好的质量。通过合成语音的听音测试,验证了该方法的有效性。索引术语:语音合成,韵律特征,F0轮廓
A total scheme of generating prosodic features from a text input was constructed. The method consists of corpus-based prediction of pauses, phone durations and fundamental frequencies (F0's), in this order, and information predicted in an earlier process is utilized in the following processes. Since prediction of F0's is done on the command values of F0 contour generation process model instead of direct F0 values, a stable and flexible control of F0 contours is possible. By adding constraints on the accent command timings as a post processing, a better quality was realized when speech was synthesized using prosodic features generated by the method. Validity of the developed method was confirmed through the listening test of the synthetic speech. Index Terms: speech synthesis, prosodic features, F0 contour