MetaSRA: normalized human sample-specific metadata for the Sequence Read Archive

MetaSRA: normalized human sample-specific metadata for the Sequence Read Archive
复制标题

DOI:
10.1093/bioinformatics/btx334
复制
发表时间:
2017-09-15
期刊:
影响因子:
5.8
通讯作者:
Dewey, Colin N.
Dewey, Colin N.
中科院分区:
生物学3区
文献类型:
--
作者:
Bernstein, Matthew N.;Doan, Anhai;Dewey, Colin N.

文献摘要

被引文献

相似文献

动机:如果能对NCBI的序列读取档案(SRA)进行综合分析,将会对生物学有很大的帮助;然而,这些数据在很大程度上仍未得到充分利用,部分原因是与每个样本相关的元数据结构不佳。向SRA提交的规则没有规定一套标准化的术语,这些术语应用于描述获得测序数据的生物样品。因此,元数据包括许多同义词、拼写变体和对外部信息源的引用。此外,由于存档中的样本数量众多,数据的手动注释仍然难以处理。由于这些原因,很难进行大规模分析,研究SRA中存在的各种疾病、组织和细胞类型的生物分子过程与表型之间的关系。结果:我们提出了MetaSRA,这是一个规范化的SRA人类样本特定元数据数据库,其模式受到ENCODE项目元数据组织的启发。该模式包括将样本映射到生物医学本体中的术语,用样本类型类别标记每个样本,并提取实值属性。我们通过一种新的计算管道自动化了这些任务。可用性和实现:MetaSRA可以在metasra.biostat.wis.edu上通过可搜索的web界面和批量下载获得。软件实现我们的计算管道可在http://github.com/deweylab/metasra-pipelineContact:cdewey@biostat.wisc.eduSupplementary信息:补充数据可在Bioinformatics在线。
Motivation: The NCBI's Sequence Read Archive (SRA) promises great biological insight if one could analyze the data in the aggregate; however, the data remain largely underutilized, in part, due to the poor structure of the metadata associated with each sample. The rules governing submissions to the SRA do not dictate a standardized set of terms that should be used to describe the biological samples from which the sequencing data are derived. As a result, the metadata include many synonyms, spelling variants and references to outside sources of information. Furthermore, manual annotation of the data remains intractable due to the large number of samples in the archive. For these reasons, it has been difficult to perform large-scale analyses that study the relationships between biomolecular processes and phenotype across diverse diseases, tissues and cell types present in the SRA.Results: We present MetaSRA, a database of normalized SRA human sample-specific metadata following a schema inspired by the metadata organization of the ENCODE project. This schema involves mapping samples to terms in biomedical ontologies, labeling each sample with a sampletype category, and extracting real-valued properties. We automated these tasks via a novel computational pipeline.Availability and implementation: The MetaSRA is available at metasra.biostat.wisc.edu via both a searchable web interface and bulk downloads. Software implementing our computational pipeline is available at http://github.com/deweylab/metasra-pipelineContact:cdewey@biostat.wisc.eduSupplementary information: Supplementary data are available at Bioinformatics online.