PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge

PARADE: A New Dataset for Paraphrase Identification Requiring Computer Science Domain Knowledge
复制标题

DOI:
10.18653/v1/2020.emnlp-main.611
复制
发表时间:
2020-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Yun He;Zhuoer Wang;Yin Zhang;Ruihong Huang;James Caverlee
Yun He;Zhuoer Wang;Yin Zhang;Ruihong Huang;James Caverlee
中科院分区:
其他
文献类型:
--
作者:
Yun He;Zhuoer Wang;Yin Zhang;Ruihong Huang;James Caverlee

文献摘要

相似文献

我们提出了一个新的基准数据集称为PARADE的释义识别,需要专门的领域知识。PARADE包含在词汇和句法层面上重叠很少但基于计算机科学领域知识在语义上等价的释义,以及在词汇和句法层面上重叠很大但基于该领域知识在语义上不等价的非释义。实验表明,最先进的神经模型和非专家的人类注释器在PARADE上的性能都很差。例如,BERT在微调后的F1得分为0.709,远低于其在其他释义识别数据集上的性能。PARADE可以作为对测试包含领域知识的模型感兴趣的研究人员的资源。我们免费提供我们的数据和代码。
We present a new benchmark dataset called PARADE for paraphrase identification that requires specialized domain knowledge. PARADE contains paraphrases that overlap very little at the lexical and syntactic level but are semantically equivalent based on computer science domain knowledge, as well as non-paraphrases that overlap greatly at the lexical and syntactic level but are not semantically equivalent based on this domain knowledge. Experiments show that both state-of-the-art neural models and non-expert human annotators have poor performance on PARADE. For example, BERT after fine-tuning achieves an F1 score of 0.709, which is much lower than its performance on other paraphrase identification datasets. PARADE can serve as a resource for researchers interested in testing models that incorporate domain knowledge. We make our data and code freely available.