Creating a Dataset for Named Entity Recognition in the Archaeology Domain

Creating a Dataset for Named Entity Recognition in the Archaeology Domain
复制标题

DOI:
--
复制
发表时间:
2020-05
期刊:
--
影响因子:
--
通讯作者:
Alex Brandsen;S. Verberne;M. Wansleeben;K. Lambers
Alex Brandsen;S. Verberne;M. Wansleeben;K. Lambers
中科院分区:
其他
文献类型:
--
作者:
Alex Brandsen;S. Verberne;M. Wansleeben;K. Lambers

文献摘要

被引文献

相似文献

在本文中,我们提出了一个训练数据集的开发荷兰命名实体识别(NER)在考古领域。该数据集的创建是因为考古学中迫切需要语义搜索,以便让考古学家在荷兰挖掘报告中找到结构化信息,目前总计约60,000(6.58亿字)并迅速增长。为了指导这项搜索任务,需要NER。我们在迭代过程中创建了严格的注释指南,然后指导五名考古学学生注释一些文件。生成的数据集包含六种实体类型(文物,时间段,地点,上下文,物种和材料)之间的约31k注释。注释者间的一致性为0.95,当我们将这些数据用于机器学习时,我们观察到F1分数从0.51增加到0.70,与在先前工作中创建的数据集上训练的机器学习模型相比。这表明数据是高质量的,并且可以自信地用于训练NER分类器。
In this paper, we present the development of a training dataset for Dutch Named Entity Recognition (NER) in the archaeology domain. This dataset was created as there is a dire need for semantic search within archaeology, in order to allow archaeologists to find structured information in collections of Dutch excavation reports, currently totalling around 60,000 (658 million words) and growing rapidly. To guide this search task, NER is needed. We created rigorous annotation guidelines in an iterative process, then instructed five archaeology students to annotate a number of documents. The resulting dataset contains ~31k annotations between six entity types (artefact, time period, place, context, species & material). The inter-annotator agreement is 0.95, and when we used this data for machine learning, we observed an increase in F1 score from 0.51 to 0.70 in comparison to a machine learning model trained on a dataset created in prior work. This indicates that the data is of high quality, and can confidently be used to train NER classifiers.