Evaluating BERT-based scientific relation classifiers for scholarly knowledge graph construction on digital library collections

Evaluating BERT-based scientific relation classifiers for scholarly knowledge graph construction on digital library collections
复制标题

DOI:
10.1007/s00799-021-00313-y
复制
发表时间:
2021-11
影响因子:
1.5
通讯作者:
Ming Jiang;J. D’Souza;S. Auer;J. Downie
Ming Jiang;J. D’Souza;S. Auer;J. Downie
中科院分区:
--
文献类型:
--
作者:
Ming Jiang;J. D’Souza;S. Auer;J. Downie

文献摘要

被引文献

相似文献

研究出版物的快速增长对数字图书馆提出了先进的信息管理技术要求。为了满足这些需求,人们提倡基于知识图结构的技术。在这种基于图的管道中,推断相关科学概念之间的语义关系是至关重要的一步。近年来,基于bert的预训练模型被广泛用于自动关系分类。尽管取得了重大进展,但其中大多数是在不同的情况下进行评估的,这限制了它们的可比性。此外,现有的方法主要是在干净的文本上进行评估,而忽略了早期学术出版物在机器扫描和光学字符识别(OCR)方面的数字化背景。在这种情况下,文本可能包含OCR噪声,从而对现有分类器的性能产生不确定性。为了解决这些限制,我们首先基于三个干净的语料库创建了ocr噪声文本。基于这些平行语料库,我们通过三个因素对八个基于bert的分类模型进行了全面的实证评估:(1)bertvariant;(2)分类策略;(3) OCR噪声影响。在干净数据上的实验表明,特定领域的预训练beredts是识别科学关系的最佳变体。一般来说,每次预测一个关系的策略优于同时识别多个关系的策略。在有噪声的语料库上,最优分类器的f分数可能会下降10%到20%左右。本研究中讨论的见解可以帮助DL利益相关者选择构建最佳基于知识图的系统的技术。
The rapid growth of research publications has placed great demands on digital libraries (DL) for advanced information management technologies. To cater to these demands, techniques relying on knowledge-graph structures are being advocated. In such graph-based pipelines, inferring semantic relations between related scientific concepts is a crucial step. Recently, BERT-based pre-trained models have been popularly explored for automatic relation classification. Despite significant progress, most of them were evaluated in different scenarios, which limits their comparability. Furthermore, existing methods are primarily evaluated on clean texts, which ignores the digitization context of early scholarly publications in terms of machine scanning and optical character recognition (OCR). In such cases, the texts may contain OCR noise, in turn creating uncertainty about existing classifiers’ performances. To address these limitations, we started by creating OCR-noisy texts based on three clean corpora. Given these parallel corpora, we conducted a thorough empirical evaluation of eightBert-based classification models by focusing on three factors: (1)Bertvariants; (2) classification strategies; and, (3) OCR noise impacts. Experiments on clean data show that the domain-specific pre-trainedBertis the best variant to identify scientific relations. The strategy of predicting a single relation each time outperforms the one simultaneously identifying multiple relations in general. The optimal classifier’s performance can decline by around 10% to 20% in F-score on the noisy corpora. Insights discussed in this study can help DL stakeholders select techniques for building optimal knowledge-graph-based systems.