Identifiers for the 21st century: How to design, provision, and reuse persistent identifiers to maximize utility and impact of life science data.

Identifiers for the 21st century: How to design, provision, and reuse persistent identifiers to maximize utility and impact of life science data.
复制标题

DOI:
10.1371/journal.pbio.2001414
复制
发表时间:
2017-06
期刊:
影响因子:
9.8
通讯作者:
Parkinson H
Parkinson H
中科院分区:
生物学1区
文献类型:
--
作者:
McMurry JA;Juty N;Blomberg N;Burdett T;Conlin T;Conte N;Courtot M;Deck J;Dumontier M;Fellows DK;Gonzalez-Beltran A;Gormanns P;Grethe J;Hastings J;Hériché JK;Hermjakob H;Ison JC;Jimenez RC;Jupp S;Kunze J;Laibe C;Le Novère N;Malone J;Martin MJ;McEntyre JR;Morris C;Muilu J;Müller W;Rocca-Serra P;Sansone SA;Sariyar M;Snoep JL;Soiland-Reyes S;Stanford NJ;Swainston N;Washington N;Williams AR;Wimalaratne SM;Winfree LM;Wolstencroft K;Goble C;Mungall CJ;Haendel MA;Parkinson H

文献摘要

被引文献

相似文献

在许多学科中,数据在数千个在线数据库(存储库,注册表和知识库)中进行了高度分散。从此类数据库中的扭动价值取决于数据科学的学科以及使整合成为可能的不起眼的砖头和砂浆。标识符是此集成基础架构的核心组成部分。利用我们的经验和其他小组的工作,我们概述了10堂课,我们已经了解了促进大规模数据集成的标识符质量和最佳实践。具体而言,我们提出了标识符从业人员(数据库提供者)应采取的措施,并应采取标识符的设计,提供和再利用。我们还概述了在各种情况下(包括作者和数据生成器)的参考标识符的重要考虑因素。尽管每个课程的重要性和相关性会因上下文而有所不同,但需要提高人们对如何避免和管理常见标识符问题的认识,尤其是与持久性和网络访问性/共振性有关的问题。我们专注于生命科学中的基于网络的标识符;但是,这些原则与其他学科广泛相关。
In many disciplines, data are highly decentralized across thousands of online databases (repositories, registries, and knowledgebases). Wringing value from such databases depends on the discipline of data science and on the humble bricks and mortar that make integration possible; identifiers are a core component of this integration infrastructure. Drawing on our experience and on work by other groups, we outline 10 lessons we have learned about the identifier qualities and best practices that facilitate large-scale data integration. Specifically, we propose actions that identifier practitioners (database providers) should take in the design, provision and reuse of identifiers. We also outline the important considerations for those referencing identifiers in various circumstances, including by authors and data generators. While the importance and relevance of each lesson will vary by context, there is a need for increased awareness about how to avoid and manage common identifier problems, especially those related to persistence and web-accessibility/resolvability. We focus strongly on web-based identifiers in the life sciences; however, the principles are broadly relevant to other disciplines.