Implementation of low-latency electrolaryngeal speech enhancement based on multi-task CLDNN

Implementation of low-latency electrolaryngeal speech enhancement based on multi-task CLDNN
复制标题

DOI:
10.23919/eusipco47968.2020.9287721
复制
发表时间:
2021-01
期刊:
2020 28th European Signal Processing Conference (EUSIPCO)
影响因子:
--
通讯作者:
Kazuhiro Kobayashi;T. Toda
Kazuhiro Kobayashi;T. Toda
中科院分区:
其他
文献类型:
--
作者:
Kazuhiro Kobayashi;T. Toda

文献摘要

相似文献

本文提出了一种基于多任务CLDNN的低延迟喉电语音增强技术。虽然EL语音可以产生相对可理解的语音,但由于机械激励信号的作用,喉切除术总是会导致语音自然度的质量下降。为了解决这一问题,提出了一种基于CLDNN的EL语音增强技术,该技术由卷积层、递归层和全连通层组成。在该技术中,基于为每个特征优化的专家CLDNN,将EL语音的输入特征向量转换为多个声码器参数,例如激励参数和频谱参数。然而,利用语音通信是困难的,因为其双向递归层导致等待话语结束的较大延迟。针对这一问题,本文提出了一种具有单向递归层的多任务CLDNN,用于低延迟EL语音增强。此外,为了获得与双向CLDNN相当的性能,我们还提出了以下技术:1)知识提取,2)数据扩充,3)语音正则化。实验结果表明,该方法能够获得与双向CLDNN相似的客观结果,并且在噪声环境下优于自然度和语音清晰度。
In this paper, we propose a low-latency speech enhancement technique for electrolaryngeal (EL) speech based on multi-task CLDNN. Although the EL speech can generate relatively intelligible speech, laryngectomees always suffer quality degradation of speech naturalness due to the mechanical excitation signals. To solve this problem, an EL speech enhancement technique based on CLDNN consisting of convolution, recurrent, and fully connected layers has been proposed. In this technique, an input feature vector of the EL speech is converted into several vocoder parameters such as excitation parameters and spectral parameters based on expert CLDNNs optimized for each feature. However, it is difficult to utilize speech communication because its bi-directional recurrent layers cause a large delay to wait for the end of the utterance. To address this issue, in this paper, we propose multi-task CLDNN with uni-directional recurrent layers for the low-latency EL speech enhancement. Moreover, to achieve comparable performance to the bi-directional CLDNN, we also propose the following techniques: 1) knowledge distillation, 2) data augmentation, and 3) phonetic regularization. The experimental results demonstrate that the proposed method makes it possible to achieve comparable objective results to the bi-directional CLDNN and outperform naturalness and speech intelligibility in the noisy condition.