Datastorr: a workflow and package for delivering successive versions of 'evolving data' directly into R

Datastorr: a workflow and package for delivering successive versions of 'evolving data' directly into R
复制标题

Datastorr:用于将“不断变化的数据”的连续版本直接交付到 R 中的工作流程和包

DOI:
--
复制
发表时间:
2019
期刊:
影响因子:
9.2
通讯作者:
W. Cornwell
W. Cornwell
中科院分区:
生物学2区
文献类型:
--
作者:
D. Falster;R. FitzJohn;Matthew W. Pennell;W. Cornwell

文献摘要

被引文献

相似文献

数据的共享和再利用已经成为现代科学的基石。多个平台现在允许轻松发布数据集。然而,到目前为止,数据共享平台在分发和与不断发展的数据集交互方面提供的功能有限——这些数据集随着时间的推移而不断增长,因为添加了更多的记录,修复了错误,创建了新的数据结构。在本文中,我们描述了一个工作流,用于维护和分发不断发展的数据集的连续版本,允许用户直接检索和加载不同的版本到R平台。我们的工作流程利用了用于开发和分发开源软件程序的连续版本的工具和平台,包括版本控制、GitHub和语义版本控制,并将这些应用于开发开源数据集的连续版本的类似过程。此外,我们认为该模型允许单个研究小组免费实现数据传递的动态和版本控制模型。
Abstract The sharing and re-use of data has become a cornerstone of modern science. Multiple platforms now allow easy publication of datasets. So far, however, platforms for data sharing offer limited functions for distributing and interacting with evolving datasets— those that continue to grow with time as more records are added, errors fixed, and new data structures are created. In this article, we describe a workflow for maintaining and distributing successive versions of an evolving dataset, allowing users to retrieve and load different versions directly into the R platform. Our workflow utilizes tools and platforms used for development and distribution of successive versions of an open source software program, including version control, GitHub, and semantic versioning, and applies these to the analogous process of developing successive versions of an open source dataset. Moreover, we argue that this model allows for individual research groups to achieve a dynamic and versioned model of data delivery at no cost.