Speeding Up Neural Machine Translation Decoding by Cube Pruning

Speeding Up Neural Machine Translation Decoding by Cube Pruning
复制标题

DOI:
10.18653/v1/d18-1460
复制
发表时间:
2018-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Wen Zhang;Liang Huang;Yang Feng;Lei Shen;Qun Liu
Wen Zhang;Liang Huang;Yang Feng;Lei Shen;Qun Liu
中科院分区:
其他
文献类型:
--
作者:
Wen Zhang;Liang Huang;Yang Feng;Lei Shen;Qun Liu

文献摘要

相似文献

尽管神经机器翻译已取得了令人瞩目的成果,但它存在翻译速度慢的问题。其直接后果是,不得不在翻译质量和速度之间进行权衡,因此其性能无法得到充分发挥。我们将一种常用于加速动态规划的技术——立方剪枝,应用于神经机器翻译以加快翻译速度。为构建等价类,相似的目标隐藏状态被合并,这使得目标端的循环神经网络(RNN)扩展操作减少,并且在庞大的目标词汇表上的softmax操作也相应减少。实验表明,在翻译质量相同甚至更好的情况下,与朴素束搜索相比,我们的方法在GPU上的翻译速度可提高3.3倍,在CPU上可提高3.5倍。
Although neural machine translation has achieved promising results, it suffers from slow translation speed. The direct consequence is that a trade-off has to be made between translation quality and speed, thus its performance can not come into full play. We apply cube pruning, a popular technique to speed up dynamic programming, into neural machine translation to speed up the translation. To construct the equivalence class, similar target hidden states are combined, leading to less RNN expansion operations on the target side and less softmax operations over the large target vocabulary. The experiments show that, at the same or even better translation quality, our method can translate faster compared with naive beam search by 3.3x on GPUs and 3.5x on CPUs.