A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice Conversion

A Preliminary Study of a Two-Stage Paradigm for Preserving Speaker Identity in Dysarthric Voice Conversion
复制标题

DOI:
10.21437/interspeech.2021-208
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Wen-Chin Huang;Kazuhiro Kobayashi;Yu-Huai Peng;Ching-Feng Liu;Yu Tsao;Hsin-Min Wang;T. Toda
Wen-Chin Huang;Kazuhiro Kobayashi;Yu-Huai Peng;Ching-Feng Liu;Yu Tsao;Hsin-Min Wang;T. Toda
中科院分区:
其他
文献类型:
--
作者:
Wen-Chin Huang;Kazuhiro Kobayashi;Yu-Huai Peng;Ching-Feng Liu;Yu Tsao;Hsin-Min Wang;T. Toda

文献摘要

相似文献

我们提出了一个新的范式来维持说话人的身份在困难语音转换(DVC)。通过统计VC可以大大改善构音障碍语音的质量,但由于构音障碍患者的正常语音几乎不可能收集,以往的工作未能恢复患者的个性。鉴于此,我们建议采用一种新颖的两阶段方法治疗DVC,该方法高度灵活,不需要患者的正常言语。首先,一个强大的并行序列到序列模型将输入的困难语音转换为参考说话者的正常语音作为中间产品,然后用变分自编码器实现的非并行、逐帧VC模型将参考语音的说话者身份转换回患者的身份,同时假定能够保持增强的质量。我们研究了几种设计方案。实验评估结果表明,我们的方法在保持说话人身份的同时提高了困难语音的质量。
We propose a new paradigm for maintaining speaker identity in dysarthric voice conversion (DVC). The poor quality of dysarthric speech can be greatly improved by statistical VC, but as the normal speech utterances of a dysarthria patient are nearly impossible to collect, previous work failed to recover the individuality of the patient. In light of this, we suggest a novel, two-stage approach for DVC, which is highly flexible in that no normal speech of the patient is required. First, a powerful parallel sequence-to-sequence model converts the input dysarthric speech into a normal speech of a reference speaker as an intermediate product, and a nonparallel, frame-wise VC model realized with a variational autoencoder then converts the speaker identity of the reference speech back to that of the patient while assumed to be capable of preserving the enhanced quality. We investigate several design options. Experimental evaluation results demonstrate the potential of our approach to improving the quality of the dysarthric speech while maintaining the speaker identity.