PUblications Metadata Augmentation (PUMA) pipeline.

PUblications Metadata Augmentation (PUMA) pipeline.
复制标题

DOI:
10.12688/f1000research.25484.2
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Burton TWY
Burton TWY
中科院分区:
其他
文献类型:
--
作者:
Butters OW;Wilson RC;Garner H;Burton TWY

文献摘要

相似文献

队列研究在很长一段时间内--通常是在参与者的生命周期内--收集、生成和分发数据。这些研究通常在其网站上发布一份出版物清单(可达数千份),以展示研究的影响,并方便搜索研究数据对其做出贡献的现有研究。在不同的研究中,搜索和探索这些出版物列表的能力差异很大。我们认为,研究出版物缺乏丰富的搜索和探索功能是进入研究数据的新用户或潜在用户的障碍,因为在特定领域可能很难找到和评估以前的工作。这些出版物清单通常也是人工整理的,导致缺乏丰富的元数据可供分析,这使得文献计量分析变得困难。我们在这里提供了一个软件管道,它聚合了来自各种第三方提供商的元数据,以支持基于Web的出版物列表搜索和探索工具。除了核心出版物元数据(即作者列表、关键字等),我们还包括第一作者的地理编码和正在进行的引文计数。这使得可以根据作者的共同位置、关键字的频率、引文概况等来描述整个研究的特征。这种丰富的出版物元数据可用于生成研究影响度量和基于网络的图形以供公开传播。此外,该管道还为文献计量学分析或科学社会研究提供研究数据集。我们使用以前发布的队列研究的出版物列表作为样本输入数据集,以显示管道的输出和效用。
Cohort studies collect, generate and distribute data over long periods of time – often over the lifecourse of their participants. It is common for these studies to host a list of publications (which can number many thousands) on their website to demonstrate the impact of the study and facilitate the search of existing research to which the study data has contributed. The ability to search and explore these publication lists varies greatly between studies. We believe a lack of rich search and exploration functionality of study publications is a barrier to entry for new or prospective users of a study’s data, since it may be difficult to find and evaluate previous work in a given area. These lists of publications are also typically manually curated, resulting in a lack of rich metadata to analyse, making bibliometric analysis difficult. We present here a software pipeline that aggregates metadata from a variety of third-party providers to power a web based search and exploration tool for lists of publications. Alongside core publication metadata (i.e. author lists, keywords etc.), we include geocoding of first authors and citation counts in our pipeline. This allows a characterisation of a study as a whole based on common locations of authors, frequency of keywords, citation profile etc. This enriched publications metadata can be useful for generating study impact metrics and web-based graphics for public dissemination. In addition, the pipeline produces a research data set for bibliometric analysis or social studies of science. We use a previously published list of publications from a cohort study as an exemplar input data set to show the output and utility of the pipeline here.