Subset Retrieval Nearest Neighbor Machine Translation

Subset Retrieval Nearest Neighbor Machine Translation
复制标题

DOI:
10.18653/v1/2023.acl-long.10
复制
发表时间:
2023
期刊:
--
影响因子:
--
通讯作者:
Hiroyuki Deguchi;Taro Watanabe;Yusuke Matsui;M. Utiyama;Hideki Tanaka;E. Sumita
Hiroyuki Deguchi;Taro Watanabe;Yusuke Matsui;M. Utiyama;Hideki Tanaka;E. Sumita
中科院分区:
其他
文献类型:
--
作者:
Hiroyuki Deguchi;Taro Watanabe;Yusuke Matsui;M. Utiyama;Hideki Tanaka;E. Sumita

文献摘要

被引文献

相似文献

k-最近邻机器翻译(kNN-MT)(Khandelwal等人,2021)通过将示例搜索纳入解码算法,提高了经过训练的神经机器翻译(NMT)模型的翻译性能。然而,解码是非常耗时的,即,比标准NMT慢大约100到1,000倍,因为在每个时间步中从并行数据的所有目标令牌中检索邻居令牌。在本文中,我们提出了“子集kNN-MT”,它通过两种方法提高了kNN-MT的解码速度:(1)从输入句子的相邻句子集合的子集中检索相邻目标标记,而不是从所有句子中检索,以及(2)有效的距离计算技术,适用于使用查找表进行子集邻居搜索。在WMT'19 De-En翻译任务和De-En和En-Ja中的域适应任务中,与kNN-MT相比,我们提出的方法实现了高达132.2倍的速度和高达1.6的BLEU分数的提高。
k-nearest-neighbor machine translation (kNN-MT) (Khandelwal et al., 2021) boosts the translation performance of trained neural machine translation (NMT) models by incorporating example-search into the decoding algorithm. However, decoding is seriously time-consuming, i.e., roughly 100 to 1,000 times slower than standard NMT, because neighbor tokens are retrieved from all target tokens of parallel data in each timestep. In this paper, we propose “Subset kNN-MT”, which improves the decoding speed of kNN-MT by two methods: (1) retrieving neighbor target tokens from a subset that is the set of neighbor sentences of the input sentence, not from all sentences, and (2) efficient distance computation technique that is suitable for subset neighbor search using a look-up table. Our proposed method achieved a speed-up of up to 132.2 times and an improvement in BLEU score of up to 1.6 compared with kNN-MT in the WMT’19 De-En translation task and the domain adaptation tasks in De-En and En-Ja.