Statistical analysis of bilingual speaker's speech for cross-language voice conversion.

Statistical analysis of bilingual speaker's speech for cross-language voice conversion.
复制标题

用于跨语言语音转换的双语说话者语音统计分析。

DOI:
10.1121/1.402284
复制
发表时间:
1991
期刊:
The Journal of the Acoustical Society of America
影响因子:
--
通讯作者:
H. Kuwabara
H. Kuwabara
中科院分区:
--
文献类型:
--
作者:
M. Abe;K. Shikano;H. Kuwabara

文献摘要

被引文献

相似文献

跨语言语音转换的目标是在翻译一个说话者的语音并用于合成另一种语言的语音时保留该说话者的语音特征。本文报告了两项初步研究,即不同语言频谱差异的统计分析和跨语言语音转换的首次尝试。对双语说话者发出的语音进行分析,以检查英语和日语之间的频谱差异。实验结果是:(1)英语和日语混合语音的码本大小应该几乎是英语或日语的码本大小的两倍; (2) 尽管许多代码向量出现在英语和日语中,但有些代码向量有在一种语言或另一种语言中占主导地位的趋势; (3)主要出现在英语中的码向量包含在音素/r/、/ae/、/f/、/s/中,主要出现在日语中的码向量包含在/i/、/u/、/N/中; (4)从听力测试来看,听众无法可靠地区分日语码本解码的英语语音和英语码本解码的英语语音之间的区别。基于码本映射的语音转换算法应用于跨语言语音转换,其性能略逊于同语言语音转换。
The goal of cross-language voice conversion is to preserve the speech characteristics of one speaker when that speaker's speech is translated and used to synthesize speech in another language. In this paper, two preliminary studies, i.e., a statistical analysis of spectrum differences in different languages and the first attempt at a cross-language voice conversion, are reported. Speech uttered by a bilingual speaker is analyzed to examine spectrum difference between English and Japanese. Experimental results are (1) the codebook size for mixed speech from English and Japanese should be almost twice the codebook size of either English or Japanese; (2) although many code vectors occurred in both English and Japanese, some have a tendency to predominate in one language or the other; (3) code vectors that predominantly occurred in English are contained in the phonemes /r/, /ae/, /f/, /s/, and code vectors that predominantly occurred in Japanese are contained in /i/, /u/, /N/; and (4) judged from listening tests, listeners cannot reliably indicate the distinction between English speech decoded by a Japanese codebook and English speech decoded by an English codebook. A voice conversion algorithm based on codebook mapping was applied to cross-language voice conversion, and its performance was somewhat less effective than for voice conversion in the same language.