Phoneme-guided Dysarthric Speech Conversion With Non-parallel Data by Joint Training

Phoneme-guided Dysarthric Speech Conversion With Non-parallel Data by Joint Training
复制标题

通过联合训练使用非并行数据进行音素引导的构音障碍语音转换

DOI:
10.1007/s11760-021-02119-6
复制
发表时间:
2022
期刊:
Signal, Image and Video Processing
影响因子:
--
通讯作者:
Tetsuya Takiguchi
Tetsuya Takiguchi
中科院分区:
--
文献类型:
--
作者:
Xunquan Chen;Atsuki Oshiro;Jinhui Chen;Ryoichi Takashima;Tetsuya Takiguchi

文献摘要

相似文献

构音障碍言语的语音结构比正常言语的语音结构更难辨别。因此,在本文中,我们提出了一个新的语音转换框架构音障碍语音学习解开音频转录表示。这种方法的新奇在于它同时将音频及其相应的转录作为训练输入。我们将从音频输入中提取的语言表示限制为接近来自转录输入的语言表示,迫使它们共享相同的分布。此外,所提出的模型可以生成适当的语言表示在测试阶段没有任何成绩单。客观和主观评价的结果表明,该方法具有更高的可懂度和更好的说话人相似度的转换语音比基线的方法。
The phonetic structures of dysarthric speech are more difficult to discriminate than those of normal speech. Therefore, in this paper, we propose a novel voice conversion framework for dysarthric speech by learning disentangled audio-transcription representations. The novelty of this method is that it simultaneously takes both audio and its corresponding transcription as training inputs. We constrain the extracted linguistic representation from the audio input to be close to the linguistic representation from the transcription input, forcing them to share the same distribution. Furthermore, the proposed model can generate appropriate linguistic representations without any transcripts during the testing stage. The results of objective and subjective evaluations showed that the proposed method exhibits higher intelligibility and better speaker similarity of the converted speech than those of the baseline approaches.