课题基金 / 基金详情

项目摘要

项目成果

RICK L. STEVENS的其他基金

相似基金

相关文献

中文摘要
翻译
这个子项目是许多研究子项目中利用 资源由NIH/NCRR资助的中心拨款提供。子项目和 调查员(PI)可能从NIH的另一个来源获得了主要资金, 并因此可以在其他清晰的条目中表示。列出的机构是 该中心不一定是调查人员的机构。 由于大量用于收集实验数据的新的高通量技术的出现,生物学领域正在经历从数据贫乏领域到数据丰富领域的根本性转变:鸟枪测序、焦磷酸测序、微阵列、芯片、Biolog表型阵列、微流控设备和流式细胞仪。今天,模拟在许多领域都难以跟上数据收集的步伐,这一点在基因组测序与基因组规模建模中表现得尤为明显。虽然在过去的十年里已经对800多个原核生物基因组进行了测序,但只发表了30个基因组规模的新陈代谢模型,而且基因组测序的速度继续加快,可能会扩大这一本已巨大的差距。今天,基因组规模的新陈代谢模型社区的范例是,它需要一年或更长时间的手工工作来产生一个新的微生物模型。然而,近年来出现了一些技术,使得自动化或加快基因组规模重建过程的各个步骤成为可能,我们最近将这些技术结合在一起,进入了自动化的基因组规模模型重建管道。虽然这条管道可以在一到五天内建造一个模型,但在这个重建过程中需要大量的计算。我们建议使用TerraGrid中的计算资源来应用这种重建过程,为每个基因组完全测序的原核生物建立新的基因组规模的代谢模型。然后,我们计划在一些高影响力的科学研究中使用这些模型,包括:(1)模拟每个代谢基因的敲除,以研究这些生物的代谢网络的稳健性,并确定未来抗菌药物开发的新的潜在靶点;(2)模拟每个微生物在各种化学条件下的生长,以确定各种微生物群落能够生存的环境;(3)预测培养每个建模生物所需的最低确定的培养条件;以及(4)模拟每个建模生物的工程,以各种可再生原材料生产具有工业价值的有机化合物。虽然这些研究产生了非常不同的结果,并吸引了根本不同的应用领域,但它们都可以通过将通量平衡分析方法应用于我们将构建的基因组规模的代谢模型来完成。我们将在拟议的基因组规模模型的重建和分析中使用的主要算法是通量平衡分析。该算法涉及的最重要的计算是线性或混合整数线性优化问题的求解。幸运的是,有许多开源软件可用于解决线性和混合整数线性优化问题。我们将应用GLPK、SCIP和BCP解算器以及我们自己定制的支持MPI的FBA软件来执行所有拟议的重建和分析计算。我们预计,在这个项目的第一年,总共需要解决104个不同的混合整数线性优化问题和1012个不同的线性优化问题,总共需要150万个CPU小时。在接下来的两年里,我们预计将需要相同数量的模拟,因为发布了更多的已测序生物体,并更新了现有生物体的注释。
英文摘要
This subproject is one of many research subprojects utilizing the resources provided by a Center grant funded by NIH/NCRR. The subproject and investigator (PI) may have received primary funding from another NIH source, and thus could be represented in other CRISP entries. The institution listed is for the Center, which is not necessarily the institution for the investigator. The field of biology is undergoing a fundamental shift from a data poor field to a data rich field thanks to the advent of numerous new high-throughput technologies for the collection of experimental data: Shotgun sequencing, Pyrosequencing, Microarrays, ChIP-chip, Biolog phenotyping arrays, microfluidic devices, and flow cytometry. Today simulation is struggling to keep pace with data collection in many areas, and nowhere is this more evident than in genome-sequencing versus genome-scale modeling. While over 800 prokaryotic genomes have been sequenced in the past ten years, only 30 genome-scale metabolic models have been published, and the pace of genome-sequencing continues to increase threatening to extend this already massive gap. Today, the paradigm of the genome-scale metabolic modeling community is that it requires a year or more of manual effort to produce a new model of a microorganism. However, technologies have emerged in recent years that make it possible to automate or expedite various steps of the genome-scale reconstruction process, and we have recently tied these technologies together into an automated genome-scale model reconstruction pipeline. While this pipeline makes it possible to construct a single model in one to five days, extensive computation is required in this reconstruction process. We are proposing to use the computational resources in the TerraGrid to apply this reconstruction process to build new genome-scale metabolic models for every prokaryote with a completely sequenced genome. We then plan to use these models in a number of high-impact scientific studies including: (1) simulating the knockout of every metabolic gene to study the robustness of the metabolic networks of these organisms and identify new potential targets for future antibacterial drug development, (2) simulating growth of each microorganism in a variety of chemical conditions to identify the environments in which various communities of microorganisms are capable of surviving, (3) predicting the minimal defined media conditions that are required in order to culture each modeled organism, and (4) simulating the engineering of each modeled organism to produce organic compounds of industrial value from a variety of renewable raw materials. While these studies produce very different results and appeal to fundamentally different application areas, they all can be accomplished by applying the Flux Balance Analysis method to the genome-scale metabolic models we will be constructing. The primary algorithm we will be using in the proposed reconstruction and analysis of genome-scale models is flux balance analysis. The most significant computation involved in this algorithm is the solving of a linear or mixed integer linear optimization problem. Fortunately, numerous open source software is available for solving linear and mixed integer linear optimization problems. We will be applying the GLPK, SCIP, and BCP solvers along with our own custom built MPI-ready FBA software to perform all of the proposed reconstruction and analysis calculations. In total, we expect that 104 distinct mixed integer linear optimization problems and 1012 distinct linear optimization problems will need to be solved in the first year of this project, requiring a total of 1.5 million CPU hours. In the two following years, we anticipate an equal number of simulations will be required due to the release of additional sequenced organism and the update of the annotations in the existing organisms.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
APPLICATION OF HIGH-PERFORMANCE COMPUTING TO THE RECONSTRUCTION, ANALYSIS, AND
  • 批准号:
    8364313
  • 项目类别:
  • 资助金额:
    $0.11万
  • 财政年份:
    2011
  • 负责人:
    RICK L. STEVENS
  • 依托单位:
Microbial Informatics Resource Core
  • 批准号:
    7700361
  • 项目类别:
  • 资助金额:
    $12.64万
  • 财政年份:
    2008
  • 负责人:
    RICK L. STEVENS
  • 依托单位:
LARGE-SCALE MOLECULAR PHYLOGENY AND COMPUTATIONAL EVIDENCE FOR HORIZONTAL GENE
  • 批准号:
    7601345
  • 项目类别:
  • 资助金额:
    $0.03万
  • 财政年份:
    2007
  • 负责人:
    RICK L. STEVENS
  • 依托单位:
LARGE-SCALE MOLECULAR PHYLOGENY AND COMPUTATIONAL EVIDENCE FOR HORIZONTAL GENE
海外基金