A Natural Language Processing Pipeline for Detecting Informal Data References in Academic Literature

A Natural Language Processing Pipeline for Detecting Informal Data References in Academic Literature
复制标题

用于检测学术文献中非正式数据引用的自然语言处理管道

DOI:
10.1002/pra2.614
复制
发表时间:
2022
影响因子:
--
通讯作者:
Hemphill, Libby
Hemphill, Libby
中科院分区:
--
文献类型:
--
作者:
Lafia, Sara;Fan, Lizhou;Hemphill, Libby

文献摘要

参考文献

被引文献

相似文献

发现出版物和它们使用的数据集之间的权威链接可能是一个劳动密集型的过程。我们引入了一个自然语言处理管道,用于检索和审查出版物,以获得对研究数据集的非正式参考,这是对数据库员工作的补充。我们首先描述管道的组成部分,然后应用它来扩展权威书目,将数千项社会科学研究与使用这些研究的数据相关出版物联系起来。这一流程增加了对文献的检索,以供审查,以便纳入与数据有关的出版物集合,并使大规模发现非正式数据参考成为可能。我们贡献了(1)一个新的命名实体识别(NER)模型,它可靠地检测非正式数据引用,以及(2)一个将社会科学文献中的条目与它们引用的数据集连接起来的数据集。总而言之,这些贡献使未来在数据参考、数据引用网络和数据重用方面的工作成为可能。
Discovering authoritative links between publications and the datasets that they use can be a labor‐intensive process. We introduce a natural language processing pipeline that retrieves and reviews publications for informal references to research datasets, which complements the work of data librarians. We first describe the components of the pipeline and then apply it to expand an authoritative bibliography linking thousands of social science studies to the data‐related publications in which they are used. The pipeline increases recall for literature to review for inclusion in data‐related collections of publications and makes it possible to detect informal data references at scale. We contribute (1) a novel Named Entity Recognition (NER) model that reliably detects informal data references and (2) a dataset connecting items from social science literature with datasets they reference. Together, these contributions enable future work on data reference, data citation networks, and data reuse.
DOI: --
发表时间: 2017
影响因子: 1
作者:
J. Hellerstein;Vikram Sreekanti;Joseph E. Gonzalez;J. Dalton;Akon Dey;Sreyashi Nag;Krishna Ramachandran;Sudhanshu Arora;A. Bhattacharyya;Shirshanka Das;Mark Donsky;Gabriel Fierro;Chang She;Carl Steinbach;V. Subramanian;Eric Sun
通讯作者: J. Hellerstein;Vikram Sreekanti;Joseph E. Gonzalez;J. Dalton;Akon Dey;Sreyashi Nag;Krishna Ramachandran;Sudhanshu Arora;A. Bhattacharyya;Shirshanka Das;Mark Donsky;Gabriel Fierro;Chang She;Carl Steinbach;V. Subramanian;Eric Sun
引用位置:改变出版物引用 Dryad 数字存储库中原始数据的做法
DOI: 10.5281/zenodo.32412
发表时间: 2016
影响因子: --
作者:
C. Mayo;Elizabeth A. Hull
通讯作者: Elizabeth A. Hull
图书馆员在环:一种用于检测学术文献中研究数据的非正式提及的自然语言处理范例
DOI: 10.48550/arxiv.2203.05112
发表时间: 2022
期刊: ArXiv
影响因子: --
作者:
Lizhou Fan;Sara Lafia;David A. Bleckley;E. Moss;A. Thomer;Libby Hemphill
通讯作者: Libby Hemphill
深度数据挖掘的新工具
DOI: 10.1029/2017eo082377
发表时间: 2017
期刊: Eos
影响因子: --
作者:
S. Peters;Ian Ross;John Czaplewski;Aimee D. Glassel;J. Husson;V. Syverson;A. Zaffos;M. Livny
通讯作者: M. Livny
识别出版物中对数据集的引用
DOI: 10.1007/978-3-642-33290-6_17
发表时间: 2012
期刊: Science, Technology, & Human Values
影响因子: --
作者:
K. Boland;Dominique Ritze;K. Eckert;Brigitte Mathiak
通讯作者: Brigitte Mathiak