Supporting data for "Datastorr: a workflow and package for delivering successive versions of 'evolving data' directly into R"
收藏资源简介:
The sharing and re-use of data has become a cornerstone of modern science. Multiple platforms now allow quick and easy data sharing. So far, however, data publishing offers limited functions for interacting with evolving datasets - those that continue to grow with time as more records are added, errors fixed, and new data structures are created. In this article, we describe a workflow for maintaining and distributing successive versions of an evolving dataset, allowing users to retrieve and load different versions directly into the R platform. Our workflow utilises tools and platforms used for development and distribution of successive versions of a open source software, including version control, GitHub and semantic versioning, and applies these to the analogous process of developing successive versions of an open source dataset. Moreover, we argue that this model allows for individual research groups to achieve a dynamic and versioned model of data delivery at no cost.



