HMM-BASED EXPRESSIVE SPEECH SYNTHESIS — TOWARDS TTS WITH ARBITRARY SPEAKING STYLES AND EMOTIONS

HMM-BASED EXPRESSIVE SPEECH SYNTHESIS — TOWARDS TTS WITH ARBITRARY SPEAKING STYLES AND EMOTIONS
复制标题

DOI:
--
复制
发表时间:
2003
期刊:
--
影响因子:
--
通讯作者:
J. Yamagishi;T. Masuko;Takao Kobayashi
J. Yamagishi;T. Masuko;Takao Kobayashi
中科院分区:
其他
文献类型:
--
作者:
J. Yamagishi;T. Masuko;Takao Kobayashi

文献摘要

被引文献

相似文献

This paper describes recent progress in our approach to generating expressive speech. A goal of text-to-speech (TTS) synthesis is to have an ability to generate natural sounding speech with arbitrary speaker’s voice characteristics, speaking styles and emotional expressions. To change voice and speaking style and/or emotion of the synthetic speech arbitrarily with maintaining its naturalness, it is required that prosodic features as well as spectral features are controlled properly. Since prosodic features are more or less related to spectral features, it is desirable to control these features simultaneously taking account of the relationship between spectrum and prosody. To resolve this problem, we have proposed several key ideas which include speaking style interpolation and adaptation for HMM-based speech synthesis. This paper focuses on these ideas and provides an overview of our approach. Moreover we show experimental results which show the effectiveness of the approach.