Decoupling Segmental and Prosodic Cues of Non-native Speech through Vector Quantization

Decoupling Segmental and Prosodic Cues of Non-native Speech through Vector Quantization
复制标题

DOI:
10.21437/interspeech.2023-2202
复制
发表时间:
2023-08
期刊:
--
影响因子:
--
通讯作者:
Waris Quamer;Anurag Das;R. Gutierrez-Osuna
Waris Quamer;Anurag Das;R. Gutierrez-Osuna
中科院分区:
其他
文献类型:
--
作者:
Waris Quamer;Anurag Das;R. Gutierrez-Osuna

文献摘要

相似文献

口音转换 (AC) 旨在将非母语人士的话语转变为母语人士的话语。与通常将口音和语音质量视为一体的语音转换相比,AC 提供了更细粒度的语音分解。本文提出了一种 AC 系统,该系统将口音进一步分解为其片段和韵律特征,并提供对两个通道的独立控制。该系统使用传统模块(声学模型、说话者/韵律编码器、seq2seq 模型)来生成重音转换,该转换结合了(1)源话语的分段特征、(2)目标话语的语音特征和(3)参考话语的韵律。然而,这种想法的天真应用会阻止系统学习和传输韵律。我们表明,矢量量化和删除重复码字使系统能够传输韵律并提高语音相似性,这一点已通过客观和感知测量得到验证。
Accent conversion (AC) seeks to transform utterances from a non-native speaker to appear native-like. Compared to voice conversion, which generally treats accent and voice quality as one, AC provides a finer-grained decomposition of speech. This paper presents an AC system that further decomposes an accent into its segmental and prosodic characteristics, and provides independent control of both channels. The system uses conventional modules (acoustic model, speaker/prosody encoders, seq2seq model) to generate accent conversions that combine (1) the segmental characteristics from a source utterance, (2) the voice characteristics from a target utterance, and (3) the prosody of a reference utterance. However, naive application of this idea prevents the system from learning and transferring prosody. We show that vector quantization and removal of repeated code-words allows the system to transfer prosody and improve voice similarity, as verified by objective and perceptual measures.