Toward hidden Markov model‐based spontaneous speech synthesis

Toward hidden Markov model‐based spontaneous speech synthesis
复制标题

基于隐马尔可夫模型的自发语音合成

DOI:
10.1121/1.4787189
复制
发表时间:
2006
影响因子:
2.4
通讯作者:
S. Furui
S. Furui
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
T. Akagawa;K. Iwano;S. Furui

文献摘要

被引文献

相似文献

本文对自发语音合成的几个问题进行了研究。虽然最先进的合成系统可以获得高度可理解的语音,但它们的自然度仍然很低。因此,要达到合成自然的、自发的语音的目标,还有很多工作要做。为了利用有限的数据量对自然语音进行建模,我们使用了一个基于HMM的语音合成器,该语音合成器基于三个特征:由HMM建模的倒谱特征,以及使用量化理论I类建模的持续时间和基频特征。模型是用从自发日语语料库(CSJ)提取的大约17min的自发演讲语音来训练的。为了进行比较,同一演讲者的发言,即朗读同一演讲的抄本,被用来训练阅读语音的类似模型。通过主观配对比较测试来评价合成语音的自发性。对18名受试者的测试结果表明,被试对被试的偏好分数高于被试。
This paper investigates several issues of spontaneous speech synthesis. Although state‐of‐the‐art synthesis systems can achieve highly intelligible speech, their naturalness is still low. Therefore, much work must still be done to achieve the goal of synthesizing natural, spontaneous speech. To model spontaneous speech using a limited amount of data, we used an HMM‐based speech synthesizer based on three features: cepstral features modeled by HMMs, and duration and fundamental frequency features modeled using Quantification Theory Type I. The models were trained with approximately 17 min of spontaneous lecture speech, from a single speaker, which was extracted from the Corpus of Spontaneous Japanese (CSJ). For comparison, utterances by the same speaker, reading a transcription of the same lecture, were used to train analogous models for read speech. Spontaneity of the synthesized speech was evaluated by subjective pair comparison tests. Results obtained from 18 subjects showed that the preference score fo...