Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation

Source and Target Bidirectional Knowledge Distillation for End-to-end Speech Translation
复制标题

DOI:
10.18653/v1/2021.naacl-main.150
复制
发表时间:
2021-04
期刊:
--
影响因子:
--
通讯作者:
H. Inaguma;T. Kawahara;Shinji Watanabe
H. Inaguma;T. Kawahara;Shinji Watanabe
中科院分区:
其他
文献类型:
--
作者:
H. Inaguma;T. Kawahara;Shinji Watanabe

文献摘要

相似文献

提高端到端语音翻译(E2E-ST)模型性能的传统方法是通过预训练以及与自动语音识别(ASR)和神经机器翻译(NMT)任务的联合训练来利用源转录。然而,由于输入方式不同,很难成功地利用源语言文本。在这项工作中,我们重点研究来自外部基于文本的 NMT 模型的序列级知识蒸馏 (SeqKD)。为了充分利用源语言信息的潜力,我们提出了反向 SeqKD,即来自目标到源反向 NMT 模型的 SeqKD。为此,我们训练双语 E2E-ST 模型来预测释义转录,作为单个解码器的辅助任务。释义是通过反向翻译从双文本翻译生成的。我们进一步提出了双向 SeqKD,其中结合了前向和后向 NMT 模型的 SeqKD。对自回归和非自回归模型的实验评估表明,每个方向的 SeqKD 都能持续提高翻译性能,并且无论模型容量如何,效果都是互补的。
A conventional approach to improving the performance of end-to-end speech translation (E2E-ST) models is to leverage the source transcription via pre-training and joint training with automatic speech recognition (ASR) and neural machine translation (NMT) tasks. However, since the input modalities are different, it is difficult to leverage source language text successfully. In this work, we focus on sequence-level knowledge distillation (SeqKD) from external text-based NMT models. To leverage the full potential of the source language information, we propose backward SeqKD, SeqKD from a target-to-source backward NMT model. To this end, we train a bilingual E2E-ST model to predict paraphrased transcriptions as an auxiliary task with a single decoder. The paraphrases are generated from the translations in bitext via back-translation. We further propose bidirectional SeqKD in which SeqKD from both forward and backward NMT models is combined. Experimental evaluations on both autoregressive and non-autoregressive models show that SeqKD in each direction consistently improves the translation performance, and the effectiveness is complementary regardless of the model capacity.