Unified Architecture for Multichannel End-to-End Speech Recognition With Neural Beamforming

Unified Architecture for Multichannel End-to-End Speech Recognition With Neural Beamforming
复制标题

DOI:
10.1109/jstsp.2017.2764276
复制
发表时间:
2017-10
影响因子:
7.5
通讯作者:
Tsubasa Ochiai;Shinji Watanabe;Takaaki Hori;J. Hershey;Xiong Xiao
Tsubasa Ochiai;Shinji Watanabe;Takaaki Hori;J. Hershey;Xiong Xiao
中科院分区:
工程技术1区
文献类型:
--
作者:
Tsubasa Ochiai;Shinji Watanabe;Takaaki Hori;J. Hershey;Xiong Xiao

文献摘要

相似文献

本文提出了一个统一的架构,端到端的自动语音识别(ASR),包括麦克风阵列信号处理,如一个国家的最先进的神经波束形成器的端到端的框架。最近,端到端ASR范式作为具有深度神经网络和隐马尔可夫模型的传统混合范式的替代方案引起了极大的研究兴趣。使用这种新的范例,我们通过将声学,语音和语言模型等ASR组件与单个神经网络集成来简化ASR架构,并优化端到端ASR目标的整体组件:生成正确的标签序列。虽然大多数现有的端到端框架主要集中在清洁环境中的ASR,但我们的目标是在嘈杂的环境中构建更现实的端到端系统。为了处理此类具有挑战性的嘈杂ASR任务,我们研究了多通道端到端ASR架构,该架构通过语音增强直接将多通道语音信号转换为文本。这种架构允许语音增强和ASR组件被联合优化,以改善端到端ASR目标,并导致在存在强背景噪声的情况下工作良好的端到端框架。我们阐述了我们所提出的方法在噪声环境中的多通道ASR基准(CHiME-4和AMI)的有效性。实验结果表明,我们提出的多通道端到端系统在具有来自延迟和求和波束形成器的增强输入的情况下获得了优于传统端到端基线的性能增益(即,BeamformIT)在字符错误率方面。此外,进一步的分析表明,我们的神经波束形成器,这是优化的端到端ASR目标,成功地学习了噪声抑制功能。
This paper proposes a unified architecture for end-to-end automatic speech recognition (ASR) to encompass microphone-array signal processing such as a state-of-the-art neural beamformer within the end-to-end framework. Recently, the end-to-end ASR paradigm has attracted great research interest as an alternative to conventional hybrid paradigms with deep neural networks and hidden Markov models. Using this novel paradigm, we simplify ASR architecture by integrating such ASR components as acoustic, phonetic, and language models with a single neural network and optimize the overall components for the end-to-end ASR objective: generating a correct label sequence. Although most existing end-to-end frameworks have mainly focused on ASR in clean environments, our aim is to build more realistic end-to-end systems in noisy environments. To handle such challenging noisy ASR tasks, we study multichannel end-to-end ASR architecture, which directly converts multichannel speech signal to text through speech enhancement. This architecture allows speech enhancement and ASR components to be jointly optimized to improve the end-to-end ASR objective and leads to an end-to-end framework that works well in the presence of strong background noise. We elaborate the effectiveness of our proposed method on the multichannel ASR benchmarks in noisy environments (CHiME-4 and AMI). The experimental results show that our proposed multichannel end-to-end system obtained performance gains over the conventional end-to-end baseline with enhanced inputs from a delay-and-sum beamformer (i.e., BeamformIT) in terms of character error rate. In addition, further analysis shows that our neural beamformer, which is optimized only with the end-to-end ASR objective, successfully learned a noise suppression function.