Nearest Neighbor Non-autoregressive Text Generation

Nearest Neighbor Non-autoregressive Text Generation
复制标题

DOI:
10.48550/arxiv.2208.12496
复制
发表时间:
2022-08
期刊:
J. Inf. Process.
影响因子:
--
通讯作者:
Ayana Niwa;Sho Takase;Naoaki Okazaki
Ayana Niwa;Sho Takase;Naoaki Okazaki
中科院分区:
其他
文献类型:
--
作者:
Ayana Niwa;Sho Takase;Naoaki Okazaki

文献摘要

被引文献

相似文献

非自回归(NAR)模型可以用比自回归模型更少的计算量生成句子,但会牺牲生成质量。先前的研究通过迭代解码解决了这个问题。本研究建议使用最近邻作为 NAR 解码器的初始状态并迭代编辑它们。我们提出了一种新颖的训练策略来学习邻居的编辑操作,以改进 NAR 文本生成。实验结果表明,所提出的方法(NeighborEdit)在 JRC-Acquis En-De 数据集(使用最近邻的机器翻译的常见基准数据集)上以更少的解码迭代(迭代次数减少了十八分之一)实现了更高的翻译质量(比普通 Transformer 高 1.69 个点)。我们还确认了所提出的方法在数据到文本任务(WikiBio)上的有效性。此外,所提出的方法在 WMT'14 En-De 数据集上的性能优于 NAR 基线。我们还报告了对所提出的方法中使用的邻居示例的分析。
Non-autoregressive (NAR) models can generate sentences with less computation than autoregressive models but sacrifice generation quality. Previous studies addressed this issue through iterative decoding. This study proposes using nearest neighbors as the initial state of an NAR decoder and editing them iteratively. We present a novel training strategy to learn the edit operations on neighbors to improve NAR text generation. Experimental results show that the proposed method (NeighborEdit) achieves higher translation quality (1.69 points higher than the vanilla Transformer) with fewer decoding iterations (one-eighteenth fewer iterations) on the JRC-Acquis En-De dataset, the common benchmark dataset for machine translation using nearest neighbors. We also confirm the effectiveness of the proposed method on a data-to-text task (WikiBio). In addition, the proposed method outperforms an NAR baseline on the WMT'14 En-De dataset. We also report analysis on neighbor examples used in the proposed method.