Cascaded Models with Cyclic Feedback for Direct Speech Translation

Cascaded Models with Cyclic Feedback for Direct Speech Translation
复制标题

DOI:
10.1109/icassp39728.2021.9413719
复制
发表时间:
2020-10
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Tsz Kin Lam;Shigehiko Schamoni;S. Riezler
Tsz Kin Lam;Shigehiko Schamoni;S. Riezler
中科院分区:
其他
文献类型:
--
作者:
Tsz Kin Lam;Shigehiko Schamoni;S. Riezler

文献摘要

被引文献

相似文献

直接语音翻译描述了仅语音输入和相应翻译可用的场景。众所周知,此类数据非常有限。我们提出了一种技术,允许自动语音识别 (ASR) 和机器翻译 (MT) 级联,以利用域内直接语音翻译数据以及域外 MT 和 ASR 数据。在预训练 MT 和 ASR 后,我们使用反馈循环,其中 MT 系统的下游性能被用作信号,通过自我训练来改进 ASR 系统,并且 MT 组件在多个 ASR 输出上进行微调,使其对拼写变化更加宽容。与使用相同架构和相同数据的组件进行端到端语音翻译的比较显示,对于德语到英语的语音翻译,LibriVoxDeEn 上的 BLEU 值提高了高达 3.8 点,而 CoVoST 上的 BLEU 值提高了 5.1 点。
Direct speech translation describes a scenario where only speech inputs and corresponding translations are available. Such data are notoriously limited. We present a technique that allows cascades of automatic speech recognition (ASR) and machine translation (MT) to exploit in-domain direct speech translation data in addition to out-of-domain MT and ASR data. After pre-training MT and ASR, we use a feed-back cycle where the downstream performance of the MT system is used as a signal to improve the ASR system by self-training, and the MT component is fine-tuned on multiple ASR outputs, making it more tolerant towards spelling variations. A comparison to end-to-end speech translation using components of identical architecture and the same data shows gains of up to 3.8 BLEU points on LibriVoxDeEn and up to 5.1 BLEU points on CoVoST for German-to-English speech translation.