Re-translation versus Streaming for Simultaneous Translation

Re-translation versus Streaming for Simultaneous Translation
复制标题

重翻译与流媒体同声翻译

DOI:
--
复制
发表时间:
2020
期刊:
International Workshop on Spoken Language Translation
影响因子:
--
通讯作者:
George F. Foster
George F. Foster
中科院分区:
--
文献类型:
--
作者:
N. Arivazhagan;Colin Cherry;Wolfgang Macherey;George F. Foster

文献摘要

被引文献

相似文献

在改进流式机器翻译方面已经取得了很大的进展,这是一种同步范例,当更多的源内容变得可用时,系统会附加一个不断增长的假设。我们研究了一个相关的问题,其中的假设超出严格的附加词的修订是允许的。这适用于诸如为音频源提供实时字幕的应用。在这种情况下,我们将自定义流方法与重新翻译进行比较,这是一种简单的策略,每个新的源令牌都会从头开始触发不同的翻译。我们发现重新翻译与最先进的流媒体系统一样好或更好,即使在允许很少修改的限制下运行。我们将这一成功归功于之前提出的数据增强技术,该技术将前缀对添加到训练数据中,与wait-k推断一起形成了流翻译的强大基线。我们还强调了重新翻译的能力,包装任意强大的MT系统的实验表明,从升级到其基础模型的大的改进。
There has been great progress in improving streaming machine translation, a simultaneous paradigm where the system appends to a growing hypothesis as more source content becomes available. We study a related problem in which revisions to the hypothesis beyond strictly appending words are permitted. This is suitable for applications such as live captioning an audio feed. In this setting, we compare custom streaming approaches to re-translation, a straightforward strategy where each new source token triggers a distinct translation from scratch. We find re-translation to be as good or better than state-of-the-art streaming systems, even when operating under constraints that allow very few revisions. We attribute much of this success to a previously proposed data-augmentation technique that adds prefix-pairs to the training data, which alongside wait-k inference forms a strong baseline for streaming translation. We also highlight re-translation’s ability to wrap arbitrarily powerful MT systems with an experiment showing large improvements from an upgrade to its base model.