Correction of Automatic Speech Recognition with Transformer Sequence-To-Sequence Model

Correction of Automatic Speech Recognition with Transformer Sequence-To-Sequence Model
复制标题

DOI:
10.1109/icassp40776.2020.9053051
复制
发表时间:
2019-10
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Oleksii Hrinchuk;Mariya Popova;Boris Ginsburg
Oleksii Hrinchuk;Mariya Popova;Boris Ginsburg
中科院分区:
其他
文献类型:
--
作者:
Oleksii Hrinchuk;Mariya Popova;Boris Ginsburg

文献摘要

被引文献

相似文献

在这项工作中,我们介绍了一个简单而有效的自动语音识别后处理模型。我们的模型具有基于transformer的编码器-解码器架构,该架构将声学模型输出“翻译”成语法和语义正确的文本。我们研究了正则化和优化模型的不同策略,并表明需要广泛的数据增强和预训练权重的初始化才能实现良好的性能。在LibriSpeech基准测试中,我们的方法在单词错误率方面比贪婪解码的基线声学模型有了显着的改善,特别是在评估数据集的噪声更大的dev-other和test-other部分。我们的模型在6-gram语言模型重新评分时的性能也优于基线,并接近Transformer-XL神经语言模型重新评分的性能。
In this work, we introduce a simple yet efficient post-processing model for automatic speech recognition. Our model has Transformer-based encoder-decoder architecture which "translates" acoustic model output into grammatically and semantically correct text. We investigate different strategies for regularizing and optimizing the model and show that extensive data augmentation and the initialization with pretrained weights are required to achieve good performance. On the LibriSpeech benchmark, our method demonstrates significant improvement in word error rate over the baseline acoustic model with greedy decoding, especially on much noisier dev-other and test-other portions of the evaluation dataset. Our model also outperforms baseline with 6-gram language model re-scoring and approaches the performance of re-scoring with Transformer-XL neural language model.