Improving the Diversity of Unsupervised Paraphrasing with Embedding Outputs
Improving the Diversity of Unsupervised Paraphrasing with Embedding Outputs
复制标题
DOI:
10.18653/v1/2021.mrl-1.15
复制
发表时间:
2021-10
期刊:
影响因子:
--
通讯作者:
M. Jegadeesan;Sachin Kumar;J. Wieting;Yulia Tsvetkov
中科院分区:
文献类型:
--
作者:
M. Jegadeesan;Sachin Kumar;J. Wieting;Yulia Tsvetkov
We present a novel technique for zero-shot paraphrase generation. The key contribution is an end-to-end multilingual paraphrasing model that is trained using translated parallel corpora to generate paraphrases into “meaning spaces” – replacing the final softmax layer with word embeddings. This architectural modification, plus a training procedure that incorporates an autoencoding objective, enables effective parameter sharing across languages for more fluent monolingual rewriting, and facilitates fluency and diversity in the generated outputs. Our continuous-output paraphrase generation models outperform zero-shot paraphrasing baselines when evaluated on two languages using a battery of computational metrics as well as in human assessment.