End-To-End Voice Conversion Via Cross-Modal Knowledge Distillation for Dysarthric Speech Reconstruction

End-To-End Voice Conversion Via Cross-Modal Knowledge Distillation for Dysarthric Speech Reconstruction
复制标题

通过跨模态知识蒸馏进行端到端语音转换,用于构音障碍语音重建

DOI:
10.1109/icassp40776.2020.9054596
复制
发表时间:
2020
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
H. Meng
H. Meng
中科院分区:
--
文献类型:
--
作者:
Disong Wang;Jianwei Yu;Xixin Wu;Songxiang Liu;Lifa Sun;Xunying Liu;H. Meng

文献摘要

被引文献

相似文献

由于不稳定韵律的修复和不精确发音的纠正,动态语音重建(DSR)是一项具有挑战性的任务。受基于序列到序列(Seq2seq)的文本到语音(TTS)合成和知识提取(KD)技术的成功启发,提出了一种新的端到端语音转换(VC)方法来解决重建任务。拟议的方法包含三个组成部分。首先,首先用转录的正常语音训练基于seq2seq的TTS。其次,以TTS系统的文本编码者为“教师”,通过训练语音编码者从转录的构音障碍语音中提取合适的语言表征,提出了跨模式KD的教师-学生框架。第三,通过将构音障碍语音直接映射到其正常版本,将前一分量的语音编码器与第一分量的注意和解码器(TTS)级联以执行DSR任务。实验结果表明,该方法可以生成自然度和清晰度都较高的语音,其中,对于语音清晰度较低和非常低的构音障碍说话人,重建语音和原始构音障碍语音的识别效果分别达到了35.4%和48.7%。
Dysarthric speech reconstruction (DSR) is a challenging task due to difficulties in repairing unstable prosody and correcting imprecise articulation. Inspired by the success of sequence-to-sequence (seq2seq) based text-to-speech (TTS) synthesis and knowledge distillation (KD) techniques, this paper proposes a novel end-to-end voice conversion (VC) method to tackle the reconstruction task. The proposed approach contains three components. First, a seq2seq based TTS is first trained with the transcribed normal speech. Second, with the text-encoder of this trained TTS system as "teacher", a teacher-student framework is proposed for cross-modal KD by training a speech-encoder to extract appropriate linguistic representations from the transcribed dysarthric speech. Third, the speech-encoder of the previous component is concatenated with the attention and decoder of the first component (TTS) to perform the DSR task, by directly mapping the dysarthric speech to its normal version. Experiments demonstrate that the proposed method can generate the speech with high naturalness and intelligibility, where the comparisons of human speech recognition between the reconstructed speech and the original dysarthric speech show that 35.4% and 48.7% absolute word error rate (WER) reduction can be achieved for dysarthric speakers with low and very low speech intelligibility, respectively.