Zero-Shot Foreign Accent Conversion without a Native Reference

Zero-Shot Foreign Accent Conversion without a Native Reference
复制标题

DOI:
10.21437/interspeech.2022-10664
复制
发表时间:
2022-09
期刊:
--
影响因子:
--
通讯作者:
Waris Quamer;Anurag Das;John M. Levis;E. Chukharev-Hudilainen;R. Gutierrez-Osuna
Waris Quamer;Anurag Das;John M. Levis;E. Chukharev-Hudilainen;R. Gutierrez-Osuna
中科院分区:
其他
文献类型:
--
作者:
Waris Quamer;Anurag Das;John M. Levis;E. Chukharev-Hudilainen;R. Gutierrez-Osuna

文献摘要

相似文献

以往的外国口音转换(FAC)方法要么在合成过程中需要来自本族语者(L1)的参考话语,要么是必须针对每个非本族语者(L2)单独训练的专用一对一系统。为了解决这两个问题,我们提出了一个新的FAC系统,可以直接从以前看不见的扬声器转换L2语音。该系统由两个独立的模块:翻译器和合成器,其操作的瓶颈功能来自语音后验图。训练翻译器将L2话语中的瓶颈特征映射到平行L1话语中的瓶颈特征。合成器是一个多对多系统,它将输入瓶颈特征映射到相应的Mel频谱图中,条件是嵌入来自L2扬声器。在推理过程中,这两个模块依次操作,以获取一个看不见的L2话语,并生成一个本地口音的梅尔频谱图。感知实验表明,与为每个L2扬声器建立专用模型的最先进的无参考系统(28.9%)相比,我们的系统实现了非母语重音的大幅减少(67%)。此外,80%的听众评价合成的话语具有相同的语音身份的L2扬声器。
Previous approaches for foreign accent conversion (FAC) ei-ther need a reference utterance from a native speaker (L1) during synthesis, or are dedicated one-to-one systems that must be trained separately for each non-native (L2) speaker. To address both issues, we propose a new FAC system that can transform L2 speech directly from previously unseen speakers. The system consists of two independent modules: a translator and a synthesizer, which operate on bottleneck features derived from phonetic posteriorgrams. The translator is trained to map bottleneck features in L2 utterances into those from a parallel L1 utterance. The synthesizer is a many-to-many system that maps input bottleneck features into the corresponding Mel-spectrograms, conditioned on an embedding from the L2 speaker. During inference, both modules operate in sequence to take an unseen L2 utterance and generate a native-accented Mel-spectrogram. Perceptual experiments show that our system achieves a large reduction (67%) in non-native accentedness compared to a state-of-the-art reference-free system (28.9%) that builds a dedicated model for each L2 speaker. Moreover, 80% of the listeners rated the synthesized utterances to have the same voice identity as the L2 speaker.