A Practical Two-Stage Training Strategy for Multi-Stream End-to-End Speech Recognition

A Practical Two-Stage Training Strategy for Multi-Stream End-to-End Speech Recognition
复制标题

DOI:
10.1109/icassp40776.2020.9053455
复制
发表时间:
2019-10
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Ruizhi Li;Gregory Sell;Xiaofei Wang;Shinji Watanabe;H. Hermansky
Ruizhi Li;Gregory Sell;Xiaofei Wang;Shinji Watanabe;H. Hermansky
中科院分区:
其他
文献类型:
--
作者:
Ruizhi Li;Gregory Sell;Xiaofei Wang;Shinji Watanabe;H. Hermansky

文献摘要

被引文献

相似文献

音频处理的多流模式,其中多个源被同时考虑,一直是信息融合的一个活跃的研究领域。我们以前的研究提供了一个很有前途的方向内端到端的自动语音识别,并行编码器的目标是捕捉不同的信息,然后基于注意力机制的流级融合联合收割机结合不同的意见。然而,随着流数量的增加导致编码器数量的增加,先前的方法可能需要大量的存储器和大量的并行数据用于联合训练。在这项工作中,我们提出了一个实用的两阶段的培训计划。阶段1是训练通用特征提取器(UFE),其中编码器输出是从用所有数据训练的单流模型产生的。阶段2制定了多流方案,旨在使用UFE特征和来自阶段1的预训练组件来单独训练注意力融合模块。实验已经进行了两个数据集,DIRHA和AMI,作为一个多流的情况下。与我们以前的方法相比,该策略实现了8.2- 32.4%的相对字错误率降低,同时始终优于几个传统的组合方法。
The multi-stream paradigm of audio processing, in which several sources are simultaneously considered, has been an active research area for information fusion. Our previous study offered a promising direction within end-to-end automatic speech recognition, where parallel encoders aim to capture diverse information followed by a stream-level fusion based on attention mechanisms to combine the different views. However, with an increasing number of streams resulting in an increasing number of encoders, the previous approach could require substantial memory and massive amounts of parallel data for joint training. In this work, we propose a practical two-stage training scheme. Stage-1 is to train a Universal Feature Extractor (UFE), where encoder outputs are produced from a single-stream model trained with all data. Stage-2 formulates a multi-stream scheme intending to solely train the attention fusion module using the UFE features and pretrained components from Stage-1. Experiments have been conducted on two datasets, DIRHA and AMI, as a multi-stream scenario. Compared with our previous method, this strategy achieves relative word error rate reductions of 8.2–32.4%, while consistently outperforming several conventional combination methods.