Correction of Automatic Speech Recognition with Transformer Sequence-To-Sequence Model
Correction of Automatic Speech Recognition with Transformer Sequence-To-Sequence Model
复制标题
DOI:
10.1109/icassp40776.2020.9053051
复制
发表时间:
2019-10
期刊:
影响因子:
--
通讯作者:
Oleksii Hrinchuk;Mariya Popova;Boris Ginsburg
中科院分区:
文献类型:
--
作者:
Oleksii Hrinchuk;Mariya Popova;Boris Ginsburg
In this work, we introduce a simple yet efficient post-processing model for automatic speech recognition. Our model has Transformer-based encoder-decoder architecture which "translates" acoustic model output into grammatically and semantically correct text. We investigate different strategies for regularizing and optimizing the model and show that extensive data augmentation and the initialization with pretrained weights are required to achieve good performance. On the LibriSpeech benchmark, our method demonstrates significant improvement in word error rate over the baseline acoustic model with greedy decoding, especially on much noisier dev-other and test-other portions of the evaluation dataset. Our model also outperforms baseline with 6-gram language model re-scoring and approaches the performance of re-scoring with Transformer-XL neural language model.