A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation

A Comparative Study on Non-Autoregressive Modelings for Speech-to-Text Generation
复制标题

DOI:
10.1109/asru51503.2021.9688157
复制
发表时间:
2021-10
期刊:
2021 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU)
影响因子:
--
通讯作者:
Yosuke Higuchi;Nanxin Chen;Yuya Fujita;H. Inaguma;Tatsuya Komatsu;Jaesong Lee;Jumon Nozaki;Tianzi W
Yosuke Higuchi;Nanxin Chen;Yuya Fujita;H. Inaguma;Tatsuya Komatsu;Jaesong Lee;Jumon Nozaki;Tianzi W
中科院分区:
其他
文献类型:
--
作者:
Yosuke Higuchi;Nanxin Chen;Yuya Fujita;H. Inaguma;Tatsuya Komatsu;Jaesong Lee;Jumon Nozaki;Tianzi W

文献摘要

相似文献

非自回归(NAR)模型在一个序列中同时生成多个输出,与自回归基线相比,这显著降低了推理速度,但代价是精度下降。NAR模型在实时应用中显示出巨大的潜力,越来越多的NAR模型在不同的领域中被探索,以缩小与AR模型的性能差距。在这项工作中,我们进行了各种NAR建模方法的端到端的自动语音识别(ASR)的比较研究。实验在最先进的环境中使用ESPnet进行。各种任务的结果为理解NAR ASR提供了有趣的发现,例如准确性-速度权衡和对长形式话语的鲁棒性。我们还表明,这些技术可以结合起来,以进一步改善和应用到NAR端到端的语音翻译。所有的实现都是公开的,以鼓励在NAR语音处理的进一步研究。
Non-autoregressive (NAR) models simultaneously generate multiple outputs in a sequence, which significantly reduces the inference speed at the cost of accuracy drop compared to autoregressive baselines. Showing great potential for real-time applications, an increasing number of NAR models have been explored in different fields to mitigate the performance gap against AR models. In this work, we conduct a comparative study of various NAR modeling methods for end-to-end automatic speech recognition (ASR). Experiments are performed in the state-of-the-art setting using ESPnet. The results on various tasks provide interesting findings for developing an understanding of NAR ASR, such as the accuracy-speed trade-off and robustness against long-form utterances. We also show that the techniques can be combined for further improvement and applied to NAR end-to-end speech translation. All the implementations are publicly available to encourage further research in NAR speech processing.