Development of a Software Pipeline for Sequence Data
Development of a Software Pipeline for Sequence Data
批准号:
7852966
负责人:
Toby Bloom
金额:
$62.1万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-30 至 2011-08-31
关键词:
AlgorithmsAllelesBiologicalBlood capillariesCatalogingCatalogsComputer softwareCopy Number PolymorphismCustomDNA Sequence RearrangementDataData SetDetectionDevelopmentEnvironmentExplosionGenerationsGenomeIndividualInformaticsInstitutesInternetLibrariesLocationMalignant NeoplasmsMarshalMeasuresMutationOutputPerformancePhaseProcessProductionProgramming LanguagesReadingResearch InfrastructureResearch PersonnelResourcesRunningSamplingSequence AnalysisServicesSomatic MutationSourceSystemTechnologyTestingTimeTranslatingUnited States National Institutes of HealthValidationVariantbiological researchcapillarycomputerized data processingcomputing resourcesdesigndetectorexperienceflexibilityhigh throughput analysisinstrumentnext generationpublic health relevanceresearch studytool
中文摘要
描述(由申请人提供):新一代测序技术在很短的时间内产生了惊人的数据量。当原始序列数据提交给NCBI时,目前NIH没有标准的管道可以以统一、稳健、快速和准确的方式处理大量数据,以产生进一步生物学研究所需的变体呼叫。对于大型合作项目,如1000基因组或TCGA,通过单一的、经过验证的分析管道一致地处理项目的所有数据对结果的质量至关重要。管道必须能够验证数据。重新校准错误率,合并跨多个来源和技术的每个样本的数据,对齐参考,并调用SNP和结构变体。此外,如果数据产量继续沿着目前的轨迹增长,这些管道每天将需要处理tb级的数据。目前,每个大型项目都在协调自己的管道基础设施和分析过程,或者协调通过不一致的过程产生的结果。此外,下一代技术使小型实验室仅用一两台仪器就能生成庞大的数据集成为可能。但这些实验室可能没有配备充分利用这些数据所需的IT和信息学基础设施。因此,他们需要在一些外部地点处理他们的数据,以使这些仪器的潜力成为现实。我们建议构建和部署一个由NCBI管理的大规模并行、高吞吐量的分析管道基础设施,并托管在Amazon Web Services (Amazon“云”)上。我们将进一步开发几个预配置的分析管道工作流,以在该基础结构上运行常见类型的序列分析。用户将能够修改和扩展预先配置的管道工作流程,或者使用我们提供的工具,随着新型测序分析的发展,设计和部署新的管道。这些新的管道将能够结合各种编程语言实现的分析算法,并且能够使用可用的计算资源尽可能并行运行,从而减少交付结果的时间。最后,我们将提供算法实现的目录,已经配置为在管道基础结构中运行,从中可以构造新的管道工作流。这些组件将包括质量重新校准步骤,snp检测器和indel检测算法。
英文摘要
DESCRIPTION (provided by applicant): Next-gen sequencing technologies are generating an incredible amount of data in a very short time span. While the raw sequence data is submitted to NCBI, at present there is no standard pipeline at NIH that can process this vast amount of data in a uniform, robust, fast and accurate manner to produce the variant calls needed for further biological research. For large collaborative projects, such as 1000 genomes or TCGA, it is critical to the quality of the results that all data for the project be processed consistently, through a single, validated analysis pipeline. The pipelines must be able to validate the data. recalibrate error rates, merge data for each sample across multiple sources and technologies, align to reference, and call SNP's and structural variants. Further, if increases in data production continue along current trajectories, these pipelines will need to process terabases of data per day. At present, every large project is coordinating its own pipeline infrastructure and analysis processes, or alternatively, reconciling results generated through inconsistent processes. Furthermore, next-generation technologies make it possible for small labs to generate huge datasets with only one or two instruments. But those labs are likely not equipped with the IT and informatics infrastructure needed to make full use of these data. They will therefore need to process their data at some external location to make the potential of these instruments a reality. We propose to build and deploy a massively-parallel, high-throughput analysis pipeline infrastructure to be managed by NCBI, and hosted at Amazon Web Services (the Amazon "cloud"). We will further develop several pre-configured analysis pipeline workflows to run common types of sequence analysis on that infrastructure. Users will be able to modify and extend the pre-configured pipeline workflows, or design and deploy new pipelines as new types of sequencing analyses develop, using tools we provide. Those new pipelines will be able to incorporate analysis algorithms implemented in a variety of programming languages, and will be able to use available compute resources to run as much as possible in parallel, thus reducing the time to delivery of results. Finally, we will provide a catalog of algorithm implementations, already configured to run within the pipeline infrastructure, from which new pipeline workflows can be constructed. These components will include quality recalibration steps, snp detectors, and indel detection algorithms.
PUBLIC HEALTH RELEVANCE: Next-generation sequencing technologies are generating an incredible amount of data in a very short time span, and analysis pipelines are needed to process this raw data to produce usable biological information. We propose to build and deploy a massively-parallel, high-throughput analysis pipeline infrastructure to be managed by NCBI. We will further develop several pre-configured analysis pipeline workflows to run common types of sequence analysis on that infrastructure.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Development of a Software Pipeline for Sequence Data
-
批准号:8542292
-
项目类别:
-
资助金额:$20.0万
-
财政年份:2009
-
负责人:Toby Bloom
-
依托单位:
海外基金