BiodivNERE: Gold standard corpora for named entity recognition and relation extraction in the biodiversity domain.

BiodivNERE: Gold standard corpora for named entity recognition and relation extraction in the biodiversity domain.
复制标题

DOI:
10.3897/bdj.10.e89481
复制
发表时间:
2022
影响因子:
1.3
通讯作者:
König-Ries B
König-Ries B
中科院分区:
环境科学与生态学4区
文献类型:
--
作者:
Abdelmageed N;Löffler F;Feddoul L;Algergawy A;Samuel S;Gaikwad J;Kazem A;König-Ries B

文献摘要

被引文献

相似文献

生物多样性是地球上各种生命的总称,包括进化、生态、生物和社会形态。为了保护生命的多样性和丰富性,必须监测生物多样性的现状及其随时间的变化,并了解驱动它的力量。这种需要导致了这一领域的大量著作出版。由此产生了大量的文本数据(出版物)和元数据(例如数据集描述)。为了支持对这些数据的管理和分析,计算机科学中的两种技术引起了人们的兴趣,即命名实体识别(NER)和关系提取(RE)。前者能够更好地发现和理解内容,而后者通过检测实体之间的连接来促进分析,从而允许我们得出结论并回答相关领域特定的问题。为了自动预测实体及其关系,可以使用机器/深度学习技术。这些技术的培训和评估需要标注语料库。在这篇文章中,我们提供了两个黄金标准语料库,用于命名实体识别(NER)和关系提取(RE),这些语料库是从生物多样性数据集元数据和摘要中生成的,可以作为评估基准,用于开发需要机器学习或深度学习技术的新的计算机支持工具。这些语料库由生物多样性专家手动标注和核实。此外,我们还解释了构建这些数据集的详细步骤。此外,我们还展示了用于标注此类语料库的类和关系的底层本体。
Biodiversity is the assortment of life on earth covering evolutionary, ecological, biological, and social forms. To preserve life in all its variety and richness, it is imperative to monitor the current state of biodiversity and its change over time and to understand the forces driving it. This need has resulted in numerous works being published in this field. With this, a large amount of textual data (publications) and metadata (e.g. dataset description) has been generated. To support the management and analysis of these data, two techniques from computer science are of interest, namely Named Entity Recognition (NER) and Relation Extraction (RE). While the former enables better content discovery and understanding, the latter fosters the analysis by detecting connections between entities and, thus, allows us to draw conclusions and answer relevant domain-specific questions. To automatically predict entities and their relations, machine/deep learning techniques could be used. The training and evaluation of those techniques require labelled corpora. In this paper, we present two gold-standard corpora for Named Entity Recognition (NER) and Relation Extraction (RE) generated from biodiversity datasets metadata and abstracts that can be used as evaluation benchmarks for the development of new computer-supported tools that require machine learning or deep learning techniques. These corpora are manually labelled and verified by biodiversity experts. In addition, we explain the detailed steps of constructing these datasets. Moreover, we demonstrate the underlying ontology for the classes and relations used to annotate such corpora.