BERTSeg: BERT Based Unsupervised Subword Segmentation for Neural Machine Translation

BERTSeg: BERT Based Unsupervised Subword Segmentation for Neural Machine Translation
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Haiyue Song;Raj Dabre;Zhuoyuan Mao;Chenhui Chu;S. Kurohashi
Haiyue Song;Raj Dabre;Zhuoyuan Mao;Chenhui Chu;S. Kurohashi
中科院分区:
其他
文献类型:
--
作者:
Haiyue Song;Raj Dabre;Zhuoyuan Mao;Chenhui Chu;S. Kurohashi

文献摘要

相似文献

现有的子词分割器要么是1)基于频率的,没有语义信息,要么是2)基于神经的,但在并行语料库上训练。为了解决这个问题,我们提出了BERTSeg,一个用于神经机器翻译的无监督神经子词分割器,它利用了characterBERT中单词的上下文语义嵌入,并最大限度地提高了子词分割的生成概率。此外,我们提出了一种基于生成概率的正则化方法,使BERTSeg能够为一个单词产生多个分割,以提高神经机器翻译的鲁棒性。实验结果表明,与BPE相比,在ALT,IWSLT 15 Vi->En,WMT 16 Ro->En和WMT 15 Fi->En数据集上,具有正则化的BERTSeg在9个平移方向上实现了高达8个BLEU点的改进。此外,BERTSeg是高效的,需要长达5分钟的训练。
Existing subword segmenters are either 1) frequency-based without semantics information or 2) neural-based but trained on parallel corpora. To address this, we present BERTSeg, an unsupervised neural subword segmenter for neural machine translation, which utilizes the contextualized semantic embeddings of words from characterBERT and maximizes the generation probability of subword segmentations. Furthermore, we propose a generation probability-based regularization method that enables BERTSeg to produce multiple segmentations for one word to improve the robustness of neural machine translation. Experimental results show that BERTSeg with regularization achieves up to 8 BLEU points improvement in 9 translation directions on ALT, IWSLT15 Vi->En, WMT16 Ro->En, and WMT15 Fi->En datasets compared with BPE. In addition, BERTSeg is efficient, needing up to 5 minutes for training.