End-to-End Integration of Speech Recognition, Dereverberation, Beamforming, and Self-Supervised Learning Representation

End-to-End Integration of Speech Recognition, Dereverberation, Beamforming, and Self-Supervised Learning Representation
复制标题

DOI:
10.1109/slt54892.2023.10023199
复制
发表时间:
2022-10
期刊:
2022 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
Yoshiki Masuyama;Xuankai Chang;Samuele Cornell;Shinji Watanabe;Nobutaka Ono
Yoshiki Masuyama;Xuankai Chang;Samuele Cornell;Shinji Watanabe;Nobutaka Ono
中科院分区:
其他
文献类型:
--
作者:
Yoshiki Masuyama;Xuankai Chang;Samuele Cornell;Shinji Watanabe;Nobutaka Ono

文献摘要

相似文献

自监督学习表示(SSLR)在自动语音识别(ASR)中已经证明了其显著的有效性,主要是针对干净的语音。最近的工作指出了在噪声环境下将SSLR与单通道语音增强相结合用于ASR的优势。本文通过对多路输入的处理,进一步推进了这种集成。通过在单个神经网络中集成去混响、波束形成、SSLR和ASR,我们提出了一种新的端到端架构。我们的系统在CHIME-4 6声道轨道上取得了文献报道的最佳性能,误字率(WER)为1.77%。虽然基于WavLM的强SSLR本身显示了良好的结果,但端到端的集成与加权功率最小化无失真响应波束形成器同时执行去混响和去噪,显著提高了WER。其有效性也在混响数据集上得到了验证。
Self-supervised learning representation (SSLR) has demonstrated its significant effectiveness in automatic speech recognition (ASR), mainly with clean speech. Recent work pointed out the strength of integrating SSLR with single-channel speech enhancement for ASR in noisy environments. This paper further advances this integration by dealing with multi-channel input. We propose a novel end-to-end architecture by integrating dereverberation, beamforming, SSLR, and ASR within a single neural network. Our system achieves the best performance reported in the literature on the CHiME-4 6-channel track with a word error rate (WER) of 1.77%. While the WavLM-based strong SSLR demonstrates promising results by itself, the end-to-end integration with the weighted power minimization distortionless response beamformer, which simultaneously performs dereverberation and denoising, improves WER significantly. Its effectiveness is also validated on the REVERB dataset.