A High-Quality Gold Standard for Citation-based Tasks

A High-Quality Gold Standard for Citation-based Tasks
复制标题

基于引文的任务的高质量黄金标准

DOI:
--
复制
发表时间:
2018
期刊:
--
影响因子:
--
通讯作者:
A. Jatowt
A. Jatowt
中科院分区:
--
文献类型:
--
作者:
Michael Färber;Alexander Thiemann;A. Jatowt

文献摘要

被引文献

相似文献

由于可用出版物的数量不断增加,在特定的引文背景下分析和推荐引文最近受到了广泛关注。虽然CiteSeerX等数据集已被创建用于评估此类任务的方法,但这些数据集显示出惊人的缺陷。当考虑到需要执行信息提取和实体链接以及实体解析时,这是可以理解的。在本文中,我们提出了一个新的评价数据集的引用依赖的任务的基础上arXiv.org出版物。我们的数据集的特点是,它在其提取的内容中几乎没有噪音,并且所有引用都与其正确的出版物相关联。除了逐句提供的纯内容外,引用的出版物通过全局标识符直接在文本中注释。参考出版物尽可能进一步链接到DBLP计算机科学参考书目。我们的数据集由超过1500万个句子组成,可免费用于研究目的。它可用于训练和测试基于引用的任务,例如推荐引用、确定引用的功能或重要性以及根据引用总结文档。
Analyzing and recommending citations within their specific citation contexts has recently received much attention due to the growing number of available publications. Although data sets such as CiteSeerX have been created for evaluating approaches for such tasks, those data sets exhibit striking defects. This is understandable when one considers that both information extraction and entity linking, as well as entity resolution, need to be performed. In this paper, we propose a new evaluation data set for citation-dependent tasks based on arXiv.org publications. Our data set is characterized by the fact that it exhibits almost zero noise in its extracted content and that all citations are linked to their correct publications. Besides the pure content, available on a sentence-by-sentence basis, cited publications are annotated directly in the text via global identifiers. As far as possible, referenced publications are further linked to the DBLP Computer Science Bibliography. Our data set consists of over 15 million sentences and is freely available for research purposes. It can be used for training and testing citation-based tasks, such as recommending citations, determining the functions or importance of citations, and summarizing documents based on their citations.