COPO: a metadata platform for brokering FAIR data in the life sciences

COPO: a metadata platform for brokering FAIR data in the life sciences
复制标题

DOI:
10.1101/782771
复制
发表时间:
2019-09
期刊:
bioRxiv
影响因子:
--
通讯作者:
Anthony Etuk;Felix Shaw;Alejandra N. González-Beltrán;David Johnson;Marie-Angélique Laporte;P. Rocca-Serra;E. Arnaud;M. Devare;P. Kersey;Susanna-Assunta Sansone;Robert P. Davey
Anthony Etuk;Felix Shaw;Alejandra N. González-Beltrán;David Johnson;Marie-Angélique Laporte;P. Rocca-Serra;E. Arnaud;M. Devare;P. Kersey;Susanna-Assunta Sansone;Robert P. Davey
中科院分区:
其他
文献类型:
--
作者:
Anthony Etuk;Felix Shaw;Alejandra N. González-Beltrán;David Johnson;Marie-Angélique Laporte;P. Rocca-Serra;E. Arnaud;M. Devare;P. Kersey;Susanna-Assunta Sansone;Robert P. Davey

文献摘要

相似文献

科学创新越来越依赖于数据和计算资源。今天的许多生命科学研究都涉及到生成、处理和重用异构数据集,这些数据集的规模呈指数级增长。对技术专家(数据科学家和生物信息学家)处理这些数据的需求空前高涨,但这些专家通常没有接受过良好数据管理实践方面的培训。也就是说,在过去的十年里,我们已经取得了长足的进步,资助者、出版商和研究人员自己都把开放、可互操作的数据作为开放科学哲学的关键组成部分。作为回应,对FAIR原则(数据应该是可查找的、可访问的、可互操作的和可重用的)的认可已经变得司空见惯。然而,在存储、管理、分析和传播遗留数据和新数据时,实施这些原则的技术和文化挑战仍然存在。COPO是一个计算系统,它试图通过使科学家能够使用社区认可的元数据集和词汇来描述他们的研究对象(原始或处理过的数据、出版物、样本、图像等),然后使用公共或机构存储库与更广泛的科学界共享,从而解决其中的一些挑战。COPO鼓励数据生成器在发布研究对象时遵循适当的元数据标准,使用语义术语为其添加含义并指定它们之间的关系。这使得数据消费者,无论是人还是机器,都可以发现、汇总和分析数据,否则这些数据将是私有的或不可见的。以现有标准为基础,推动科学数据传播的最新水平,同时最大限度地减少数据出版和共享的负担。COPO是完全开源的,可以在GitHub上免费获得https://github.com/collaborative-open-plant-omics。供社区使用的平台的公共实例以及更多信息可以在copo-project.org上找到。
Scientific innovation is increasingly reliant on data and computational resources. Much of today’s life science research involves generating, processing, and reusing heterogeneous datasets that are growing exponentially in size. Demand for technical experts (data scientists and bioinformaticians) to process these data is at an all-time high, but these are not typically trained in good data management practices. That said, we have come a long way in the last decade, with funders, publishers, and researchers themselves making the case for open, interoperable data as a key component of an open science philosophy. In response, recognition of the FAIR Principles (that data should be Findable, Accessible, Interoperable and Reusable) has become commonplace. However, both technical and cultural challenges for the implementation of these principles still exist when storing, managing, analysing and disseminating both legacy and new data. COPO is a computational system that attempts to address some of these challenges by enabling scientists to describe their research objects (raw or processed data, publications, samples, images, etc.) using community-sanctioned metadata sets and vocabularies, and then use public or institutional repositories to share it with the wider scientific community. COPO encourages data generators to adhere to appropriate metadata standards when publishing research objects, using semantic terms to add meaning to them and specify relationships between them. This allows data consumers, be they people or machines, to find, aggregate, and analyse data which would otherwise be private or invisible. Building upon existing standards to push the state of the art in scientific data dissemination whilst minimising the burden of data publication and sharing. Availability COPO is entirely open source and freely available on GitHub at https://github.com/collaborative-open-plant-omics. A public instance of the platform for use by the community, as well as more information, can be found at copo-project.org.