DMDD: A Large-Scale Dataset for Dataset Mentions Detection

DMDD: A Large-Scale Dataset for Dataset Mentions Detection
复制标题

DOI:
10.1162/tacl_a_00592
复制
发表时间:
2023-05
影响因子:
10.9
通讯作者:
Huitong Pan;Qi Zhang;E. Dragut;Cornelia Caragea;Longin Jan Latecki
Huitong Pan;Qi Zhang;E. Dragut;Cornelia Caragea;Longin Jan Latecki
中科院分区:
人文科学1区
文献类型:
--
作者:
Huitong Pan;Qi Zhang;E. Dragut;Cornelia Caragea;Longin Jan Latecki

文献摘要

相似文献

摘要数据集名称的识别是科学文献信息自动提取的一项关键任务,使研究人员能够理解和识别研究机会。然而,现有的数据集提及检测语料库在大小和命名多样性方面受到限制。在本文中,我们介绍了数据集提及检测数据集(DMDD),最大的公开可用的语料库这项任务。DMDD由DMDD主语料库组成,包括31,219篇科学文章,其中超过449,000个数据集以文本跨度格式弱注释,以及一个评估集,其中包括450篇出于评估目的手动注释的科学文章。我们使用DMDD来建立数据集提及检测和链接的基线性能。通过分析各种模型在DMDD上的性能,我们能够识别数据集提及检测中的开放问题。我们邀请社区使用我们的数据集作为挑战,开发新的数据集提及检测模型。
Abstract The recognition of dataset names is a critical task for automatic information extraction in scientific literature, enabling researchers to understand and identify research opportunities. However, existing corpora for dataset mention detection are limited in size and naming diversity. In this paper, we introduce the Dataset Mentions Detection Dataset (DMDD), the largest publicly available corpus for this task. DMDD consists of the DMDD main corpus, comprising 31,219 scientific articles with over 449,000 dataset mentions weakly annotated in the format of in-text spans, and an evaluation set, which comprises 450 scientific articles manually annotated for evaluation purposes. We use DMDD to establish baseline performance for dataset mention detection and linking. By analyzing the performance of various models on DMDD, we are able to identify open problems in dataset mention detection. We invite the community to use our dataset as a challenge to develop novel dataset mention detection models.