A framework for organizing cancer-related variations from existing databases, publications and NGS data using a High-performance Integrated Virtual Environment (HIVE).

A framework for organizing cancer-related variations from existing databases, publications and NGS data using a High-performance Integrated Virtual Environment (HIVE).
复制标题

DOI:
10.1093/database/bau022
复制
发表时间:
2014
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Mazumder R
Mazumder R
中科院分区:
其他
文献类型:
--
作者:
Wu TJ;Shamsaddini A;Pan Y;Smith K;Crichton DJ;Simonyan V;Mazumder R

文献摘要

参考文献

被引文献

相似文献

Years of sequence feature curation by UniProtKB/Swiss-Prot, PIR-PSD, NCBI-CDD, RefSeq and other database biocurators has led to a rich repository of information on functional sites of genes and proteins. This information along with variation-related annotation can be used to scan human short sequence reads from next-generation sequencing (NGS) pipelines for presence of non-synonymous single-nucleotide variations (nsSNVs) that affect functional sites. This and similar workflows are becoming more important because thousands of NGS data sets are being made available through projects such as The Cancer Genome Atlas (TCGA), and researchers want to evaluate their biomarkers in genomic data. BioMuta, an integrated sequence feature database, provides a framework for automated and manual curation and integration of cancer-related sequence features so that they can be used in NGS analysis pipelines. Sequence feature information in BioMuta is collected from the Catalogue of Somatic Mutations in Cancer (COSMIC), ClinVar, UniProtKB and through biocuration of information available from publications. Additionally, nsSNVs identified through automated analysis of NGS data from TCGA are also included in the database. Because of the petabytes of data and information present in NGS primary repositories, a platform HIVE (High-performance Integrated Virtual Environment) for storing, analyzing, computing and curating NGS data and associated metadata has been developed. Using HIVE, 31 979 nsSNVs were identified in TCGA-derived NGS data from breast cancer patients. All variations identified through this process are stored in a Curated Short Read archive, and the nsSNVs from the tumor samples are included in BioMuta. Currently, BioMuta has 26 cancer types with 13 896 small-scale and 308 986 large-scale study-derived variations. Integration of variation data allows identifications of novel or common nsSNVs that can be prioritized in validation studies. Database URL: BioMuta: http://hive.biochemistry.gwu.edu/tools/biomuta/index.php; CSR: http://hive.biochemistry.gwu.edu/dna.cgi?cmd=csr; HIVE: http://hive.biochemistry.gwu.edu
来自1,092个人基因组的遗传变异的综合图。
DOI: 10.1038/nature11632
发表时间: 2012-11-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1038/nprot.2009.86
发表时间: 2009-01-01
期刊: NATURE PROTOCOLS
影响因子: 14.8
作者:
Kumar, Prateek;Henikoff, Steven;Ng, Pauline C.
通讯作者: Ng, Pauline C.
DOI: 10.1093/nar/gkr854
发表时间: 2012-01
影响因子: 14.9
作者:
Kodama Y;Shumway M;Leinonen R;International Nucleotide Sequence Database Collaboration
通讯作者: International Nucleotide Sequence Database Collaboration
DOI: 10.1093/nar/gkx1095
发表时间: 2018-01-04
影响因子: 14.9
作者:
NCBI Resource Coordinators
通讯作者: NCBI Resource Coordinators
DOI: 10.1016/j.gpb.2012.10.003
发表时间: 2013-04-01
影响因子: 9.5
作者:
Karagiannis, Konstantinos;Simonyan, Vahan;Mazumder, Raja
通讯作者: Mazumder, Raja