A Compression-Based Multiple Subword Segmentation for Neural Machine Translation

A Compression-Based Multiple Subword Segmentation for Neural Machine Translation
复制标题

DOI:
10.3390/electronics11071014
复制
发表时间:
2022-04-01
期刊:
影响因子:
2.9
通讯作者:
Sakamoto, Hiroshi
Sakamoto, Hiroshi
中科院分区:
工程技术3区
文献类型:
--
作者:
Nonaka, Keita;Yamanouchi, Kazutaka;Sakamoto, Hiroshi

文献摘要

被引文献

相似文献

在本研究中,我们提出了一种简单有效的基于数据压缩算法的子词分词预处理方法。基于压缩的子词分词作为神经机器翻译训练数据的预处理方法,近年来备受关注。其中,与传统方法相比,BPE/BPE-dropout是最快、最有效的方法之一;然而,基于压缩的方法有一个缺点,即由于确定性而难以生成多个分割。为了克服这个困难,我们重点研究了一种称为局部一致解析(LCP)的随机字符串算法,该算法已被用于实现最佳压缩。利用LCP的随机解析机制,我们提出了LCP-dropout用于多子词分割,提高了BPE/BPE-dropout,并且我们表明它在从特别小的训练数据中学习时优于各种基线。
In this study, we propose a simple and effective preprocessing method for subword segmentation based on a data compression algorithm. Compression-based subword segmentation has recently attracted significant attention as a preprocessing method for training data in neural machine translation. Among them, BPE/BPE-dropout is one of the fastest and most effective methods compared to conventional approaches; however, compression-based approaches have a drawback in that generating multiple segmentations is difficult due to the determinism. To overcome this difficulty, we focus on a stochastic string algorithm, called locally consistent parsing (LCP), that has been applied to achieve optimum compression. Employing the stochastic parsing mechanism of LCP, we propose LCP-dropout for multiple subword segmentation that improves BPE/BPE-dropout, and we show that it outperforms various baselines in learning from especially small training data.