Improving Large-scale Paraphrase Acquisition and Generation

Improving Large-scale Paraphrase Acquisition and Generation
复制标题

DOI:
10.48550/arxiv.2210.03235
复制
发表时间:
2022-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yao Dou;Chao Jiang;Wei Xu
Yao Dou;Chao Jiang;Wei Xu
中科院分区:
其他
文献类型:
--
作者:
Yao Dou;Chao Jiang;Wei Xu

文献摘要

相似文献

本文讨论了现有基于Twitter的释义数据集中的质量问题,并讨论了使用两种不同的释义定义来识别和生成任务的必要性。我们在Twitter(MultiPIT)语料库中提出了一个新的多主题复述,该复述由总共13万个具有众包(MultiPIT_Crowd)和专家(MultiPIT_Expert)注释的句子对组成,使用两种不同的复述定义来识别复述,此外还包括一个多参考测试集(MultiPIT_NMR)和一个自动构建的大型训练集(MultiPIT_Auto)。通过改进的数据标注质量和特定于任务的释义定义,在我们的数据集上微调的最佳预训练语言模型达到了84.2F1的最先进性能,用于自动释义识别。此外,我们的实验结果还表明,与在Quora、MSCOCO和ParaNMT等其他语料库上微调的对应模型相比,在MultiPIT_Auto上训练的复述生成模型生成的复述更加多样化和高质量。
This paper addresses the quality issues in existing Twitter-based paraphrase datasets, and discusses the necessity of using two separate definitions of paraphrase for identification and generation tasks. We present a new Multi-Topic Paraphrase in Twitter (MultiPIT) corpus that consists of a total of 130k sentence pairs with crowdsoursing (MultiPIT_crowd) and expert (MultiPIT_expert) annotations using two different paraphrase definitions for paraphrase identification, in addition to a multi-reference test set (MultiPIT_NMR) and a large automatically constructed training set (MultiPIT_Auto) for paraphrase generation. With improved data annotation quality and task-specific paraphrase definition, the best pre-trained language model fine-tuned on our dataset achieves the state-of-the-art performance of 84.2 F1 for automatic paraphrase identification. Furthermore, our empirical results also demonstrate that the paraphrase generation models trained on MultiPIT_Auto generate more diverse and high-quality paraphrases compared to their counterparts fine-tuned on other corpora such as Quora, MSCOCO, and ParaNMT.