Development of a Software Pipeline for Sequence Data
Development of a Software Pipeline for Sequence Data
批准号:
7944084
负责人:
Steven Andrew McCarroll
金额:
$59.76万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-30 至 2013-08-31
关键词:
AlgorithmsAllelesBiologicalBlood capillariesCatalogingCatalogsComputer softwareCopy Number PolymorphismCustomDNA Sequence RearrangementDataData SetDetectionDevelopmentEnvironmentExplosionGenerationsGenomeIndividualInformaticsInstitutesInternetLibrariesLocationMalignant NeoplasmsMarshalMeasuresMutationOutputPerformancePhaseProcessProductionProgramming LanguagesReadingResearch InfrastructureResearch PersonnelResourcesRunningSamplingSequence AnalysisServicesSomatic MutationSourceSystemTechnologyTestingTimeTranslatingUnited States National Institutes of HealthValidationVariantbiological researchcapillarycomputerized data processingcomputing resourcesdesigndetectorexperienceflexibilityhigh throughput analysisinfrastructure developmentinstrumentnext generationpublic health relevanceresearch studytool
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): Next-gen sequencing technologies are generating an incredible amount of data in a very short time span. While the raw sequence data is submitted to NCBI, at present there is no standard pipeline at NIH that can process this vast amount of data in a uniform, robust, fast and accurate manner to produce the variant calls needed for further biological research. For large collaborative projects, such as 1000 genomes or TCGA, it is critical to the quality of the results that all data for the project be processed consistently, through a single, validated analysis pipeline. The pipelines must be able to validate the data. recalibrate error rates, merge data for each sample across multiple sources and technologies, align to reference, and call SNP's and structural variants. Further, if increases in data production continue along current trajectories, these pipelines will need to process terabases of data per day. At present, every large project is coordinating its own pipeline infrastructure and analysis processes, or alternatively, reconciling results generated through inconsistent processes. Furthermore, next-generation technologies make it possible for small labs to generate huge datasets with only one or two instruments. But those labs are likely not equipped with the IT and informatics infrastructure needed to make full use of these data. They will therefore need to process their data at some external location to make the potential of these instruments a reality. We propose to build and deploy a massively-parallel, high-throughput analysis pipeline infrastructure to be managed by NCBI, and hosted at Amazon Web Services (the Amazon "cloud"). We will further develop several pre-configured analysis pipeline workflows to run common types of sequence analysis on that infrastructure. Users will be able to modify and extend the pre-configured pipeline workflows, or design and deploy new pipelines as new types of sequencing analyses develop, using tools we provide. Those new pipelines will be able to incorporate analysis algorithms implemented in a variety of programming languages, and will be able to use available compute resources to run as much as possible in parallel, thus reducing the time to delivery of results. Finally, we will provide a catalog of algorithm implementations, already configured to run within the pipeline infrastructure, from which new pipeline workflows can be constructed. These components will include quality recalibration steps, snp detectors, and indel detection algorithms.
PUBLIC HEALTH RELEVANCE: Next-generation sequencing technologies are generating an incredible amount of data in a very short time span, and analysis pipelines are needed to process this raw data to produce usable biological information. We propose to build and deploy a massively-parallel, high-throughput analysis pipeline infrastructure to be managed by NCBI. We will further develop several pre-configured analysis pipeline workflows to run common types of sequence analysis on that infrastructure.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
2/3-Genetic Analysis of the International Cohort Collection for Bipolar Disorder
-
批准号:8861932
-
项目类别:
-
资助金额:$17.7万
-
财政年份:2015
-
负责人:Steven Andrew McCarroll
-
依托单位:
2/3-Genetic Analysis of the International Cohort Collection for Bipolar Disorder
-
批准号:9052837
-
项目类别:
-
资助金额:$18.3万
-
财政年份:2015
-
负责人:Steven Andrew McCarroll
-
依托单位:
2/3-Whole Genome Sequencing for Schizophrenia and Bipolar Disorder in the GPC
-
批准号:8806061
-
项目类别:
-
资助金额:$326.18万
-
财政年份:2014
-
负责人:Steven Andrew McCarroll
-
依托单位:
2/3-Whole Genome Sequencing for Schizophrenia and Bipolar Disorder in the GPC
-
批准号:8930191
-
项目类别:
-
资助金额:$326.18万
-
财政年份:2014
-
负责人:Steven Andrew McCarroll
-
依托单位:
2/3-Whole Genome Sequencing for Schizophrenia and Bipolar Disorder in the GPC
-
批准号:9306200
-
项目类别:
-
资助金额:$327.95万
-
财政年份:2014
-
负责人:Steven Andrew McCarroll
-
依托单位:
2/3-Whole Genome Sequencing for Schizophrenia and Bipolar Disorder in the GPC
-
批准号:9107509
-
项目类别:
-
资助金额:$327.95万
-
财政年份:2014
-
负责人:Steven Andrew McCarroll
-
依托单位:
Structurally complex genome loci in human populations and human phenotypes
-
批准号:10211665
-
项目类别:
-
资助金额:$72.25万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Multi-allelic copy number variation of the human genome
-
批准号:8344049
-
项目类别:
-
资助金额:$50.0万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Accurate analysis of genome structural variation using large-scale sequence data
-
批准号:8236219
-
项目类别:
-
资助金额:$44.8万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Accurate analysis of genome structural variation using large-scale sequence data
-
批准号:8416344
-
项目类别:
-
资助金额:$39.77万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Multi-allelic copy number variation of the human genome
-
批准号:8532954
-
项目类别:
-
资助金额:$47.75万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Multi-allelic copy number variation of the human genome
-
批准号:8704768
-
项目类别:
-
资助金额:$49.0万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Accurate analysis of genome structural variation using large-scale sequence data
-
批准号:8606864
-
项目类别:
-
资助金额:$40.85万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Structurally complex genome loci in human populations and human phenotypes
-
批准号:10468727
-
项目类别:
-
资助金额:$72.47万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Multi-allelic forms of human genome structural variation
-
批准号:10192865
-
项目类别:
-
资助金额:$33.9万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Structurally complex genome loci in human populations and human phenotypes
-
批准号:10686008
-
项目类别:
-
资助金额:$72.47万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Population-Based Approaches to Genome Structure and Structural Variation
-
批准号:9335937
-
项目类别:
-
资助金额:$65.54万
-
财政年份:2012
-
负责人:Steven Andrew McCarroll
-
依托单位:
Computational and Statistical Genomics Analysis Core
-
批准号:9923743
-
项目类别:
-
资助金额:$16.74万
-
财政年份:--
-
负责人:Steven Andrew McCarroll
-
依托单位:
Critical periods and complement regulation in diverse CNS cell types
-
批准号:9923739
-
项目类别:
-
资助金额:$52.81万
-
财政年份:--
-
负责人:Steven Andrew McCarroll
-
依托单位:
海外基金