A statistical model based fundamental frequency synthesizer for Mandarin speech.

A statistical model based fundamental frequency synthesizer for Mandarin speech.
复制标题

基于统计模型的普通话语音基频合成器。

DOI:
10.1121/1.404276
复制
发表时间:
1992
影响因子:
2.4
通讯作者:
S. M. Lee
S. M. Lee
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
S. Chen;S. Chang;S. M. Lee

文献摘要

被引文献

相似文献

提出了一种基于统计模型的汉语文语转换基频合成方法。具体而言,统计模型来确定F0轮廓模式的音节和语言特征表示的上下文之间的关系。该模型的参数从一个大的训练集的重复话语经验估计。语音规则通过训练过程自动推导并隐式记忆在模型中。在合成过程中,从给定的输入文本中提取上下文特征,然后使用训练好的模型通过Viterbi算法找到音节的F0轮廓模式的最佳估计。该方法可以看作是采用一种随机文法来减少合成过程中每个决策点处F0轮廓模式的候选数。虽然输入文本的各个层次上的语言特征可以被纳入模型中,但在本研究中仅使用了从相邻音节中提取的一些相关上下文特征。该方法的性能进行了检查,通过模拟使用的数据库组成的9个重复的112个相同的文本,所有的发言由一个单一的扬声器的陈述性的重复话语。通过对训练好的模型进行仔细的分析,我们发现该模型隐含了降音效果和一些连读变调规则。实验结果表明,77.56%的合成F0轮廓符合原始自然语音的VQ量化对应。通过非正式的听力测试证实了合成语音的自然性。
A novel method based on a statistical model for the fundamental-frequency (F0) synthesis in Mandarin text-to-speech is proposed. Specifically, a statistical model is employed to determine the relationship between F0 contour patterns of syllables and linguistic features representing the context. Parameters of the model were empirically estimated from a large training set of sentential utterances. Phonologic rules are then automatically deduced through the training process and implicitly memorized in the model. In the synthesis process, contextual features are extracted from a given input text, and the best estimates of F0 contour patterns of syllable are then found by a Viterbi algorithm using the well-trained model. This method can be regarded as employing a stochastic grammar to reduce the number of candidates of F0 contour pattern at each decision point of synthesis. Although linguistic features on various levels of input text can be incorporated into the model, only some relevant contextual features extracted from neighboring syllables were used in this study. Performance of this method was examined by simulation using a database composed of nine repetitions of 112 declarative sentential utterances of the same text, all spoken by a single speaker. By closely examining the well-trained model, some evidence was found to show that the declination effect as well as several sandhi rules are implicitly contained in the model. Experimental results show that 77.56% of synthesized F0 contours coincide with the VQ-quantized counterpart of the original natural speech. Naturalness of the synthesized speech was confirmed by an informal listening test.