A new method for FO tracking errors fix and generation in HMM-based Mandarin speech synthesis using generation process model

A new method for FO tracking errors fix and generation in HMM-based Mandarin speech synthesis using generation process model
复制标题

一种使用生成过程模型在基于 HMM 的普通话语音合成中修复和生成 FO 跟踪错误的新方法

DOI:
10.1109/icosp.2010.5656850
复制
发表时间:
2010
期刊:
IEEE 10th INTERNATIONAL CONFERENCE ON SIGNAL PROCESSING PROCEEDINGS
影响因子:
--
通讯作者:
N. Minematsu
N. Minematsu
中科院分区:
--
文献类型:
--
作者:
Miaomiao Wang;Miaomiao Wen;K. Hirose;N. Minematsu

文献摘要

参考文献

被引文献

相似文献

基于隐马尔可夫模型的文语转换系统可以灵活地对谱参数和韵律参数进行建模,从而产生高质量的合成语音。然而,当用于训练的特征向量是噪声时,合成语音的质量下降。在所有的噪声特征中,基音跟踪误差和相应的有缺陷的浊音/清音(VU)决定是语音质量问题的两个关键因素。这些误差也会增大音素时长的均方根误差。在基于HMM的TTS持续时间通常是统计建模使用状态持续时间概率分布和持续时间预测看不见的上下文。使用丰富的上下文特征使得合成无需高级语言知识。在本文中,一个F0生成过程模型被用来重新估计F0值的基音跟踪误差的区域,以及在清音区域。对每个普通话音素施加VU的先验知识,并将其用于VU判决。另外,我们设计了两套语法特征分别用来改善普通话音素和停顿时长的预测。
The HMM-based Text-to-Speech System can produce high quality synthetic speech with flexible modeling of spectral and prosodie parameters. However the quality of synthetic speech degrades when feature vectors used in training are noisy. Among all noisy features, pitch tracking errors and corresponding flawed voiced/unvoiced (VU) decisions are the two key factors in voice quality problems. Also these errors will enlarge the RMSE of phoneme duration. In HMM-based TTS durations are typically modeled statistically using state duration probability distributions and duration prediction for unseen contexts. Use of rich context features enables synthesis without high-level linguistic knowledge. In this paper, an F0 generation process model is used to re-estimate F0 values in the regions of pitch tracking errors, as well as in unvoiced regions. A prior knowledge of VU is imposed in each Mandarin phoneme and they are used for VU decision. Also we design two sets of syntax features to improve Mandarin phone and pause duration prediction respectively.
DOI: --
发表时间: 2004-12
期刊: IEICE Trans. Inf. Syst.
影响因子: --
作者:
D. Arifianto;Tomohiro Tanaka;T. Masuko;Takao Kobayashi
通讯作者: D. Arifianto;Tomohiro Tanaka;T. Masuko;Takao Kobayashi
DOI: 10.1109/tasl.2009.2016394
发表时间: 2009-08
期刊: IEEE Transactions on Audio, Speech, and Language Processing
影响因子: --
作者:
J. Yamagishi;Takashi Nose;H. Zen;Zhenhua Ling;T. Toda;K. Tokuda;Simon King;S. Renals
通讯作者: J. Yamagishi;Takashi Nose;H. Zen;Zhenhua Ling;T. Toda;K. Tokuda;Simon King;S. Renals
DOI: 10.1109/icassp.2000.861820
发表时间: 2000-06
期刊: 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100)
影响因子: --
作者:
K. Tokuda;Takayoshi Yoshimura;T. Masuko;Takao Kobayashi;T. Kitamura
通讯作者: K. Tokuda;Takayoshi Yoshimura;T. Masuko;Takao Kobayashi;T. Kitamura