SDCF: semi-automatically structured dataset of citation functions

SDCF: semi-automatically structured dataset of citation functions
复制标题

DOI:
10.1007/s11192-022-04471-x
复制
发表时间:
2022-07
期刊:
影响因子:
3.9
通讯作者:
Basuki Setio;Masatoshi Tsuchiya
Basuki Setio;Masatoshi Tsuchiya
中科院分区:
管理学3区
文献类型:
--
作者:
Basuki Setio;Masatoshi Tsuchiya

文献摘要

相似文献

引用函数的自动检测越来越受到人们的关注,这也是学术论文作者引用前人作品的原因。用于这样的任务的机器学习方法需要由各种引用函数标签组成的大型数据集。然而,现有的数据集包含一些实例和有限数量的标签。此外,大多数标签都是使用狭窄的研究领域建立的。针对这些问题,本文提出了一种半自动的方法来开发基于两种类型的数据集的引用函数的大型数据集。第一种类型包含5668个手动标记的实例,用于开发引用函数的新标记方案,第二种类型是自动构建的最终数据集。我们的标签方案涵盖了计算机科学各个领域的论文,产生了5个大标签和21个细粒度标签。为了验证该方案,采用两个注释器对421个实例进行注释实验,这些实例产生的Cohen Kappa值为0.85(对于粗标签)和0.71(对于细粒度标签)。在此之后,我们进行了两个分类阶段,即,过滤,细粒度的使用第一个数据集建立模型。该分类遵循几种情况,包括在低资源环境中的主动学习(AL)。实验结果表明,基于BERT的人工智能算法在滤波阶段的准确率达到了90.29%,优于其他方法。在细粒度阶段,基于SciBERT的AL策略获得了81.15%的准确率,略低于非AL策略。这些结果表明,AL是有前途的,因为它需要不到一半的数据集。考虑到标签的数量,本文发布了由1,840,815个实例组成的最大数据集。
There is increasing research interest in the automatic detection of citation functions, which is why authors of academic papers cite previous works. A machine learning approach for such a task requires a large dataset consisting of varied labels of citation functions. However, existing datasets contain a few instances and a limited number of labels. Furthermore, most labels have been built using narrow research fields. Addressing these issues, this paper proposes a semiautomatic approach to develop a large dataset of citation functions based on two types of datasets. The first type contains 5668 manually labeled instances to develop a new labeling scheme of citation functions, and the second type is the final dataset that is built automatically. Our labeling scheme covers papers from various areas of computer science, resulting in fivecoarselabels and 21fine-grainedlabels. To validate the scheme, two annotators were employed for annotation experiments on 421 instances that produced Cohen’s Kappa values of 0.85 forcoarselabels and 0.71 forfine-grainedlabels. Following this, we performed two classification stages, i.e.,filtering,andfine-grainedto build models using the first dataset. The classification followed several scenarios, including active learning (AL) in a low-resource setting. Our experiments show that Bidirectional Encoder Representations from Transformers (BERT)-based AL achieved 90.29% accuracy, which outperformed other methods in thefilteringstage. In thefine-grainedstage, the SciBERT-based AL strategy achieved a competitive 81.15% accuracy, which was slightly lower than the non-AL strategy. These results show that the AL is promising since it requires less than half of the dataset. Considering the number of labels, this paper released the largest dataset consisting of 1,840,815 instances.