Phoneme-guided Dysarthric Speech Conversion With Non-parallel Data by Joint Training
Phoneme-guided Dysarthric Speech Conversion With Non-parallel Data by Joint Training
复制标题
通过联合训练使用非并行数据进行音素引导的构音障碍语音转换
DOI:
10.1007/s11760-021-02119-6
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Tetsuya Takiguchi
中科院分区:
文献类型:
--
作者:
Xunquan Chen;Atsuki Oshiro;Jinhui Chen;Ryoichi Takashima;Tetsuya Takiguchi
The phonetic structures of dysarthric speech are more difficult to discriminate than those of normal speech. Therefore, in this paper, we propose a novel voice conversion framework for dysarthric speech by learning disentangled audio-transcription representations. The novelty of this method is that it simultaneously takes both audio and its corresponding transcription as training inputs. We constrain the extracted linguistic representation from the audio input to be close to the linguistic representation from the transcription input, forcing them to share the same distribution. Furthermore, the proposed model can generate appropriate linguistic representations without any transcripts during the testing stage. The results of objective and subjective evaluations showed that the proposed method exhibits higher intelligibility and better speaker similarity of the converted speech than those of the baseline approaches.