Multichannel End-to-end Speech Recognition

Multichannel End-to-end Speech Recognition
复制标题

DOI:
--
复制
发表时间:
2017-03
期刊:
--
影响因子:
--
通讯作者:
Tsubasa Ochiai;Shinji Watanabe;Takaaki Hori;J. Hershey
Tsubasa Ochiai;Shinji Watanabe;Takaaki Hori;J. Hershey
中科院分区:
其他
文献类型:
--
作者:
Tsubasa Ochiai;Shinji Watanabe;Takaaki Hori;J. Hershey

文献摘要

被引文献

相似文献

语音识别领域正处于范式转变之中:端到端神经网络正在挑战隐马尔可夫模型作为核心技术的主导地位。在递归编码器-解码器架构中使用注意机制解决了动态时间对齐问题,允许声学和语言建模组件的联合端到端训练。在本文中,我们扩展了端到端的框架,包括麦克风阵列信号处理的噪声抑制和语音增强的声学编码网络。这允许在识别架构内联合优化波束成形组件,以提高端到端语音识别目标。噪声语音基准测试(CHiME-4和AMI)的实验表明,我们的多通道端到端系统优于基于注意力的基线与传统的自适应波束形成器的输入。
The field of speech recognition is in the midst of a paradigm shift: end-to-end neural networks are challenging the dominance of hidden Markov models as a core technology. Using an attention mechanism in a recurrent encoder-decoder architecture solves the dynamic time alignment problem, allowing joint end-to-end training of the acoustic and language modeling components. In this paper we extend the end-to-end framework to encompass microphone array signal processing for noise suppression and speech enhancement within the acoustic encoding network. This allows the beamforming components to be optimized jointly within the recognition architecture to improve the end-to-end speech recognition objective. Experiments on the noisy speech benchmarks (CHiME-4 and AMI) show that our multichannel end-to-end system outperformed the attention-based baseline with input from a conventional adaptive beamformer.