课题基金 / 基金详情

Algorithms for the Analysis of Data from Massively-parallel Genome Sequencing

Algorithms for the Analysis of Data from Massively-parallel Genome Sequencing
大规模并行基因组测序数据分析算法
批准号:
0844494
负责人:
Mihai Pop
金额:
$37.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-04-01 至 2013-03-31

项目摘要

项目成果

Mihai Pop的其他基金

相似基金

相关文献

中文摘要
翻译
新一代DNA测序技术正在给现代生物学研究带来革命性的变化。科学家们现在可以在短短几天内用一个测序仪器生成大致相当于整个人类基因组(约30亿个碱基对的DNA)。直到最近,如此大量的数据只能在大型基因组中心使用数百台测序仪产生。这些数据的分析因其规模而变得复杂——测序仪器的单次运行产生tb级的信息,通常需要对现有计算基础设施进行大规模扩展。该项目正在开发用于分析新一代测序数据的并行算法,特别关注在谷歌和IBM支持的高度分布式计算集群上实现的Map-Reduce范式。该项目主要集中于开发序列比对和序列组装的算法。基因组数据分析中的关键任务是什么?并且涉及到字符串匹配和图算法对Map-Reduce范式的适应。这项工作将有可能导致并行化的基因组分析软件,使研究人员能够通过网络规模的计算资源分析新一代测序数据,从而避免建立和维护本地高性能计算基础设施的需要。在此项目中开发的软件将在开源许可下提供,以鼓励广泛使用并使未来的研究成为可能。该研究与研究生和本科生的教学和指导相结合,工作结果将通过期刊出版物和会议报告传播。
英文摘要
New generation DNA sequencing technologies are revolutionizing modern biological research. Scientists can now generate the rough equivalent of an entire human genome (~3 billion base-pairs of DNA) in just a few days with one single sequencing instrument. Until recently, such amounts of data could only be generated at large genome centers using hundreds of sequencers. The analysis of these data is complicated by their size - a single run of a sequencing instrument yields terabytes of information, often requiring a significant scale-up of the existing computational infrastructure. This project is developing parallel algorithms for analyzing new generation sequencing data with a specific focus on the Map-Reduce paradigm implemented on a highly-distributed computing cluster supported by Google and IBM. The project is primarily focused on developing algorithms for sequence alignment and sequence assembly ? critical tasks in the analysis of genomic data ? and involves the adaptation of string matching and graph algorithms to the Map-Reduce paradigm. This work will potentially lead to parallelism-enabled genomic analysis software that will allow researchers to analyze new generation sequencing data through web-scale computational resources, thereby obviating the need for establishing and maintaining a local high-performance computing infrastructure. The software developed during this project is being made available under an open-source license in order to encourage broad use and to enable future research. The research is integrated with teaching and mentoring of graduate and undergraduate students and the results of the work will be disseminated through journal publications and conference presentations.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
REU Site: Undergraduate Bioinformatics Research in Data Science for Genomics
III: AF: Medium: Collaborative Research: Scalable and Highly Accurate Methods for Metagenomics
III: Small: Genome Assembly Using Sparse Sequence Information
III-CXT-Small: Graphs to Diversity: extracting genomic variation from sequence graphs
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
  • 批准号:
    --
  • 项目类别:
    外国学者研究基金项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    USHARANI HAREESH GOVINDARA JAN
  • 依托单位:
基于Meta-analysis的新疆棉花灌水增产模型研究
  • 批准号:
    41601604
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    22.0万元
  • 批准年份:
    2016
  • 负责人:
    赵爱琴
  • 依托单位:
大规模微阵列数据组的meta-analysis方法研究
  • 批准号:
    31100958
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    20.0万元
  • 批准年份:
    2011
  • 负责人:
    赵洪雅
  • 依托单位: