Improving the Diversity of Unsupervised Paraphrasing with Embedding Outputs

Improving the Diversity of Unsupervised Paraphrasing with Embedding Outputs
复制标题

DOI:
10.18653/v1/2021.mrl-1.15
复制
发表时间:
2021-10
期刊:
ArXiv
影响因子:
--
通讯作者:
M. Jegadeesan;Sachin Kumar;J. Wieting;Yulia Tsvetkov
M. Jegadeesan;Sachin Kumar;J. Wieting;Yulia Tsvetkov
中科院分区:
其他
文献类型:
--
作者:
M. Jegadeesan;Sachin Kumar;J. Wieting;Yulia Tsvetkov

文献摘要

相似文献

我们提出了一种新的零弹释义生成技术。关键贡献是一个端到端的多语言释义模型,该模型使用翻译的平行语料库进行训练,以生成释义到“意义空间”-用词嵌入取代最终的softmax层。这种架构上的修改,加上一个包含自动编码目标的训练过程,可以实现跨语言的有效参数共享,从而实现更流畅的单语言重写,并促进生成输出的流畅性和多样性。我们的连续输出释义生成模型在使用一组计算指标以及人类评估的两种语言上进行评估时,优于零射击释义基线。
We present a novel technique for zero-shot paraphrase generation. The key contribution is an end-to-end multilingual paraphrasing model that is trained using translated parallel corpora to generate paraphrases into “meaning spaces” – replacing the final softmax layer with word embeddings. This architectural modification, plus a training procedure that incorporates an autoencoding objective, enables effective parameter sharing across languages for more fluent monolingual rewriting, and facilitates fluency and diversity in the generated outputs. Our continuous-output paraphrase generation models outperform zero-shot paraphrasing baselines when evaluated on two languages using a battery of computational metrics as well as in human assessment.