The Sequence Read Archive: explosive growth of sequencing data.

The Sequence Read Archive: explosive growth of sequencing data.
复制标题

DOI:
10.1093/nar/gkr854
复制
发表时间:
2012-01
影响因子:
14.9
通讯作者:
International Nucleotide Sequence Database Collaboration
International Nucleotide Sequence Database Collaboration
中科院分区:
生物学2区
文献类型:
--
作者:
Kodama Y;Shumway M;Leinonen R;International Nucleotide Sequence Database Collaboration

文献摘要

被引文献

相似文献

新一代测序平台正在以显著更高的通量和更低的成本产生数据。这种能力的一部分专门用于个人和社区科学项目。随着这些项目的出版,原始测序数据集被提交到主要的下一代序列数据存档,即序列读取存档(SRA)。可重复性实验数据是可重复性科学进步的关键。SRA是作为国际核苷酸序列数据库合作(INSDC)的一部分,作为下一代序列数据的公共存储库而建立的。INSDC由国家生物技术信息中心(NCBI)、欧洲生物信息学研究所(EBI)和日本DNA数据库(DDBJ)组成。SRA可在NCBI的www.ncbi.nlm.nih.gov/sra、EBI的www.ebi.ac.uk/ena和 DDBJ的trace.ddbj.nig.ac.jp上访问。在本文中,我们介绍了SRA的内容和结构,并报告了更新的元数据结构、提交文件格式和支持的测序平台。我们还简要概述了我们对爆炸性数据增长挑战的各种应对措施。
New generation sequencing platforms are producing data with significantly higher throughput and lower cost. A portion of this capacity is devoted to individual and community scientific projects. As these projects reach publication, raw sequencing datasets are submitted into the primary next-generation sequence data archive, the Sequence Read Archive (SRA). Archiving experimental data is the key to the progress of reproducible science. The SRA was established as a public repository for next-generation sequence data as a part of the International Nucleotide Sequence Database Collaboration (INSDC). INSDC is composed of the National Center for Biotechnology Information (NCBI), the European Bioinformatics Institute (EBI) and the DNA Data Bank of Japan (DDBJ). The SRA is accessible at www.ncbi.nlm.nih.gov/sra from NCBI, at www.ebi.ac.uk/ena from EBI and at trace.ddbj.nig.ac.jp from DDBJ. In this article, we present the content and structure of the SRA and report on updated metadata structures, submission file formats and supported sequencing platforms. We also briefly outline our various responses to the challenge of explosive data growth.