Phylesystem: a git-based data store for community-curated phylogenetic estimates.

Phylesystem: a git-based data store for community-curated phylogenetic estimates.
复制标题

DOI:
10.1093/bioinformatics/btv276
复制
发表时间:
2015-09-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Smith SA
Smith SA
中科院分区:
其他
文献类型:
--
作者:
McTavish EJ;Hinchliff CE;Allman JF;Brown JW;Cranston KA;Holder MT;Rees JA;Smith SA

文献摘要

被引文献

相似文献

动机:可以使用Dryad (Vision, 2010)或TreeBASE (Sanderson et al., 1994)等通用平台对已发表研究的系统发育估计进行存档。这些服务在确保系统发育研究的透明度和可重复性方面发挥着至关重要的作用。然而,数字树数据文件通常需要一些编辑(例如重新根)来提高系统发育陈述的准确性和可重用性。此外,在树和单一公共分类学的分类群中建立尖端标签之间的映射,极大地提高了其他研究人员重用系统发育估计的能力。由于整理已发表的系统发育评估的过程并非没有错误,因此保留对树的编辑来源的完整记录对于开放性至关重要,这使编辑能够获得工作的荣誉,并使整理过程中引入的错误更容易纠正。结果:在这里,我们报告了软件基础设施的发展,以支持生物学家社区对系统发育数据的开放管理。系统的后端通过提交到git存储库,为标准的数据库操作提供了创建、读取、更新和删除记录的接口。对树的编辑历史记录由git的版本控制功能保存。将此数据存储托管在GitHub (http://github.com/)上,使用许多开发人员熟悉的工具提供对数据存储的开放访问。我们已经部署了一个运行‘ phylessystem -api ’的服务器,它封装了git和GitHub的交互。开放生命之树项目还开发并部署了一个JavaScript应用程序,该应用程序使用phylessystem -api和其他web服务来输入和管理已发布的系统发育声明。可用性和实现:web服务层的源代码可在https://github.com/OpenTreeOfLife/phylesystem-api上获得。克隆数据存储的路径:https://github.com/OpenTreeOfLife/phylesystem。使用系统web服务的web应用程序部署在http://tree.opentreeoflife.org/curator。该工具的代码可从https://github.com/OpenTreeOfLife/opentree获得。联系:mtholder@gmail.com
Motivation: Phylogenetic estimates from published studies can be archived using general platforms like Dryad (Vision, 2010) or TreeBASE (Sanderson et al., 1994). Such services fulfill a crucial role in ensuring transparency and reproducibility in phylogenetic research. However, digital tree data files often require some editing (e.g. rerooting) to improve the accuracy and reusability of the phylogenetic statements. Furthermore, establishing the mapping between tip labels used in a tree and taxa in a single common taxonomy dramatically improves the ability of other researchers to reuse phylogenetic estimates. As the process of curating a published phylogenetic estimate is not error-free, retaining a full record of the provenance of edits to a tree is crucial for openness, allowing editors to receive credit for their work and making errors introduced during curation easier to correct. Results: Here, we report the development of software infrastructure to support the open curation of phylogenetic data by the community of biologists. The backend of the system provides an interface for the standard database operations of creating, reading, updating and deleting records by making commits to a git repository. The record of the history of edits to a tree is preserved by git’s version control features. Hosting this data store on GitHub (http://github.com/) provides open access to the data store using tools familiar to many developers. We have deployed a server running the ‘phylesystem-api’, which wraps the interactions with git and GitHub. The Open Tree of Life project has also developed and deployed a JavaScript application that uses the phylesystem-api and other web services to enable input and curation of published phylogenetic statements. Availability and implementation: Source code for the web service layer is available at https://github.com/OpenTreeOfLife/phylesystem-api. The data store can be cloned from: https://github.com/OpenTreeOfLife/phylesystem. A web application that uses the phylesystem web services is deployed at http://tree.opentreeoflife.org/curator. Code for that tool is available from https://github.com/OpenTreeOfLife/opentree. Contact: mtholder@gmail.com