The NIH Open Citation Collection: A public access, broad coverage resource

The NIH Open Citation Collection: A public access, broad coverage resource
复制标题

DOI:
10.1371/journal.pbio.3000385
复制
发表时间:
2019-10-01
期刊:
影响因子:
9.8
通讯作者:
Santangelo, George M.
Santangelo, George M.
中科院分区:
生物学1区
文献类型:
--
作者:
Hutchins, B. Ian;Baker, Kirk L.;Santangelo, George M.

文献摘要

被引文献

相似文献

引文数据仍然隐藏在专有的、限制性的许可协议之后,这增加了希望使用这些数据的分析人员的进入壁垒,增加了进行大规模分析的费用,并降低了结论的稳健性和可重复性。在过去的几年里,美国国立卫生研究院(NIH)的组合分析办公室(OPA)一直在汇总和增强可以公开共享的引文数据。在这里,我们描述了NIH开放引文收集(NIH-OCC),一个公共访问的生物医学研究数据库,免费提供给社区。这个数据集是从MedLine、PubMed Central(PMC)和CrossRef等不受限制的数据源中精心生成的,现在是NIH iCite分析平台中提供的引文统计数据的基础。我们还包括来自机器学习管道的数据,该管道可以识别,提取,解析和消除互联网上全文文章中的参考文献。公开引文链接在iCite的重大更新中向公众提供(https:icite.od.nih.gov)。
Citation data have remained hidden behind proprietary, restrictive licensing agreements, which raises barriers to entry for analysts wishing to use the data, increases the expense of performing large-scale analyses, and reduces the robustness and reproducibility of the conclusions. For the past several years, the National Institutes of Health (NIH) Office of Portfolio Analysis (OPA) has been aggregating and enhancing citation data that can be shared publicly. Here, we describe the NIH Open Citation Collection (NIH-OCC), a public access database for biomedical research that is made freely available to the community. This dataset, which has been carefully generated from unrestricted data sources such as MedLine, PubMed Central (PMC), and CrossRef, now underlies the citation statistics delivered in the NIH iCite analytic platform. We have also included data from a machine learning pipeline that identifies, extracts, resolves, and disambiguates references from full-text articles available on the internet. Open citation links are available to the public in a major update of iCite (https://icite.od.nih.gov).