SnoVault and encodeD: A novel object-based storage system and applications to ENCODE metadata

SnoVault and encodeD: A novel object-based storage system and applications to ENCODE metadata
复制标题

DOI:
10.1371/journal.pone.0175310
复制
发表时间:
2017-04-12
期刊:
影响因子:
3.7
通讯作者:
Cherry, J. Michael
Cherry, J. Michael
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hitz, Benjamin C.;Rowe, Laurence D.;Cherry, J. Michael

文献摘要

被引文献

相似文献

DNA元件百科全书(ENCODE)项目是一项持续的合作努力,旨在创建一个全面的功能元件目录,该项目在人类基因组计划完成后不久就开始了。目前的数据库超过450个细胞系和组织的6500个实验,使用广泛的实验技术来研究H. sapiens和M.肌肉基因组所有ENCODE实验数据、元数据和相关的计算分析都提交给ENCODE数据协调中心(DCC)进行验证、跟踪、存储、统一处理,并分发给社区资源和科学界。随着数据量的增加,实验细节的识别和组织变得越来越复杂,需要精心策划。ENCODE DCC创建了一个称为SnoVault的通用软件系统,该系统支持元数据和文件提交、用于元数据存储的数据库、用于显示元数据的网页以及用于查询元数据的强大API。该软件是完全开源的,代码和安装说明可以在http://github.com/ENCODE-DCC/snovault/(通用数据库)和http://github.com/ENCODE-DCC/encoded/上找到,以ENCODE的方式存储基因组数据。核心数据库引擎SnoVault(完全独立于ENCODE、基因组数据或生物信息学数据)已作为单独的Python包发布。
The Encyclopedia of DNA elements (ENCODE) project is an ongoing collaborative effort to create a comprehensive catalog of functional elements initiated shortly after the completion of the Human Genome Project. The current database exceeds 6500 experiments across more than 450 cell lines and tissues using a wide array of experimental techniques to study the chromatin structure, regulatory and transcriptional landscape of the H. sapiens and M. musculus genomes. All ENCODE experimental data, metadata, and associated computational analyses are submitted to the ENCODE Data Coordination Center (DCC) for validation, tracking, storage, unified processing, and distribution to community resources and the scientific community. As the volume of data increases, the identification and organization of experimental details becomes increasingly intricate and demands careful curation. The ENCODE DCC has created a general purpose software system, known as SnoVault, that supports metadata and file submission, a database used for metadata storage, web pages for displaying the metadata and a robust API for querying the metadata. The software is fully open-source, code and installation instructions can be found at: http://github.com/ENCODE-DCC/snovault/ (for the generic database) and http://github.com/ENCODE-DCC/encoded/ to store genomic data in the manner of ENCODE. The core database engine, SnoVault (which is completely independent of ENCODE, genomic data, or bioinformatic data) has been released as a separate Python package.