Uniform genomic data analysis in the NCI Genomic Data Commons.

Uniform genomic data analysis in the NCI Genomic Data Commons.
复制标题

DOI:
10.1038/s41467-021-21254-9
复制
发表时间:
2021-02-22
影响因子:
16.6
通讯作者:
Grossman RL
Grossman RL
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Zhang Z;Hernandez K;Savage J;Li S;Miller D;Agrawal S;Ortuno F;Staudt LM;Heath A;Grossman RL

文献摘要

参考文献

被引文献

相似文献

美国国家癌症研究所(NCI‘s)基因组数据共享中心(GDC)的目标是为癌症研究社区提供一个统一处理的基因组和相关临床数据的数据库,以支持精确医学中的数据共享和协作分析。最初的GDC数据集包括来自NCI TCGA和TARGET计划的基因组、表观基因组、蛋白质组、临床和其他数据。GDC的数据生产于2015年6月开始,使用基于OpenStack的私有云。截至2016年6月,GDC已经分析了50,000多个原始测序数据输入,以及多种其他数据类型。使用最新的人类基因组参考构建GRCh38,GDC生成了从比对读取到体细胞突变、基因表达、miRNA表达、DNA甲基化状态和拷贝数变异的各种数据类型。在本文中,我们描述了用于处理和协调GDC中的数据的管道和工作流。生成的数据以及来自TCGA和TARGET的原始输入文件可在全球数据中心数据门户和遗留档案(https://gdc.cancer.gov/).)下载和探索性分析基因组数据共享库包含来自TCGA和目标数据集的基因组、表观基因组、蛋白质组和临床数据。在这里,作者描述了如何将这些不同的数据集整合在一起的分析方法。
The goal of the National Cancer Institute’s (NCI’s) Genomic Data Commons (GDC) is to provide the cancer research community with a data repository of uniformly processed genomic and associated clinical data that enables data sharing and collaborative analysis in the support of precision medicine. The initial GDC dataset include genomic, epigenomic, proteomic, clinical and other data from the NCI TCGA and TARGET programs. Data production for the GDC started in June, 2015 using an OpenStack-based private cloud. By June of 2016, the GDC had analyzed more than 50,000 raw sequencing data inputs, as well as multiple other data types. Using the latest human genome reference build GRCh38, the GDC generated a variety of data types from aligned reads to somatic mutations, gene expression, miRNA expression, DNA methylation status, and copy number variation. In this paper, we describe the pipelines and workflows used to process and harmonize the data in the GDC. The generated data, as well as the original input files from TCGA and TARGET, are available for download and exploratory analysis at the GDC Data Portal and Legacy Archive (https://gdc.cancer.gov/). The Genomic Data Commons repository contains genomic, epigenomic, proteomic and clinical data from the TCGA and TARGET datasets. Here, the authors describe the analysis methods for how these divergent datasets were integrated together.
DOI: 10.1093/bioinformatics/btu638
发表时间: 2015-01-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Anders S;Pyl PT;Huber W
通讯作者: Huber W
DOI: 10.1186/s13059-016-0974-4
发表时间: 2016-06-06
期刊: Genome biology
影响因子: 12.3
作者:
McLaren W;Gil L;Hunt SE;Riat HS;Ritchie GR;Thormann A;Flicek P;Cunningham F
通讯作者: Cunningham F
实现癌症基因组数据的共同愿景。
DOI: 10.1056/nejmp1607591
发表时间: 2016-09-22
期刊: The New England journal of medicine
影响因子: --
作者:
Grossman RL;Heath AP;Ferretti V;Varmus HE;Lowy DR;Kibbe WA;Staudt LM
通讯作者: Staudt LM
DOI: 10.1038/nature20805
发表时间: 2017-01-12
期刊: Nature
影响因子: 64.8
作者:
Cancer Genome Atlas Research Network;Analysis Working Group: Asan University;BC Cancer Agency;Brigham and Women’s Hospital;Broad Institute;Brown University;Case Western Reserve University;Dana-Farber Cancer Institute;Duke University;Greater Poland Cancer Centre;Harvard Medical School;Institute for Systems Biology;KU Leuven;Mayo Clinic;Memorial Sloan Kettering Cancer Center;National Cancer Institute;Nationwide Children’s Hospital;Stanford University;University of Alabama;University of Michigan;University of North Carolina;University of Pittsburgh;University of Rochester;University of Southern California;University of Texas MD Anderson Cancer Center;University of Washington;Van Andel Research Institute;Vanderbilt University;Washington University;Genome Sequencing Center: Broad Institute;Washington University in St. Louis;Genome Characterization Centers: BC Cancer Agency;Broad Institute;Harvard Medical School;Sidney Kimmel Comprehensive Cancer Center at Johns Hopkins University;University of North Carolina;University of Southern California Epigenome Center;University of Texas MD Anderson Cancer Center;Van Andel Research Institute;Genome Data Analysis Centers: Broad Institute;Brown University:;Harvard Medical School;Institute for Systems Biology;Memorial Sloan Kettering Cancer Center;University of California Santa Cruz;University of Texas MD Anderson Cancer Center;Biospecimen Core Resource: International Genomics Consortium;Research Institute at Nationwide Children’s Hospital;Tissue Source Sites: Analytic Biologic Services;Asan Medical Center;Asterand Bioscience;Barretos Cancer Hospital;BioreclamationIVT;Botkin Municipal Clinic;Chonnam National University Medical School;Christiana Care Health System;Cureline;Duke University;Emory University;Erasmus University;Indiana University School of Medicine;Institute of Oncology of Moldova;International Genomics Consortium;Invidumed;Israelitisches Krankenhaus Hamburg;Keimyung University School of Medicine;Memorial Sloan Kettering Cancer Center;National Cancer Center Goyang;Ontario Tumour Bank;Peter MacCallum Cancer Centre;Pusan National University Medical School;Ribeirão Preto Medical School;St. Joseph’s Hospital &Medical Center;St. Petersburg Academic University;Tayside Tissue Bank;University of Dundee;University of Kansas Medical Center;University of Michigan;University of North Carolina at Chapel Hill;University of Pittsburgh School of Medicine;University of Texas MD Anderson Cancer Center;Disease Working Group: Duke University;Memorial Sloan Kettering Cancer Center;National Cancer Institute;University of Texas MD Anderson Cancer Center;Yonsei University College of Medicine;Data Coordination Center: CSRA Inc.;Project Team: National Institutes of Health
通讯作者: Project Team: National Institutes of Health
DOI: 10.1101/gr.135350.111
发表时间: 2012-09
期刊: Genome research
影响因子: 7
作者:
Harrow J;Frankish A;Gonzalez JM;Tapanari E;Diekhans M;Kokocinski F;Aken BL;Barrell D;Zadissa A;Searle S;Barnes I;Bignell A;Boychenko V;Hunt T;Kay M;Mukherjee G;Rajan J;Despacio-Reyes G;Saunders G;Steward C;Harte R;Lin M;Howald C;Tanzer A;Derrien T;Chrast J;Walters N;Balasubramanian S;Pei B;Tress M;Rodriguez JM;Ezkurdia I;van Baren J;Brent M;Haussler D;Kellis M;Valencia A;Reymond A;Gerstein M;Guigó R;Hubbard TJ
通讯作者: Hubbard TJ