课题基金 / 基金详情

Development of MUFOLD for Building High-Accuracy Protein Structure Models

Development of MUFOLD for Building High-Accuracy Protein Structure Models
开发用于建立高精度蛋白质结构模型的 MUFOLD
批准号:
8656715
负责人:
DONG XU
金额:
$27.89万
依托单位国家:
美国
项目类别:
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-07-01 至 2017-04-30

项目摘要

项目成果

DONG XU的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):建议项目的长期目标是提供一个全面的平台,MUFOLD,用于高效和一致准确的蛋白质三级结构预测。MUFOLD将帮助实验生物学家了解他们感兴趣的蛋白质的结构和功能,从而促进实验设计的假设。我们将把重点放在资金机会公告的第二个目标--“已知结构的远程同源的高精度模型”,其中规定“这些模型的质量应接近X射线结构或高分辨率核磁共振结构,主链原子和侧链原子的RMSD小于2埃,所有蛋白质靶标的RMSD应一致。”具体地说,我们将整合生物信息学技术、图和网络理论、计算算法、全局优化方法、统计评估等,开发一个基于模板的结构预测系统,该系统将模型生成、模型质量评估(QA)和模型精化无缝集成在一起。首先,我们将深入应用已知模板库(PDB)中的相关信息以及多层QA方法来指导在小的目标构象空间中高效地生成模型,这将有助于计算效率和QA方法选择的有限数量的模型。其次,我们将通过整合一个模型的不同QA分数及其与为同一目标蛋白质生成的其他模型的结构关系来提高QA的整体识别能力。第三,我们将开发一种基于种群的模型精化协议,该协议集成了不同级别的QA和高效的模型生成技术,以提高模型的整体质量。我们的目标是1)提高预测速度,使200~300个残基的目标蛋白质的预测在多核台式机上在几分钟内完成;2)增强从生成的候选模型中选择最佳模型的QA能力,并将目前平均~10点的GDT-TS损失从最佳可用模型降低到5点;3)实现对远距离同源蛋白质的预测精度平均在2埃RMSD以内的主链和侧链原子;以及4)与PSI(蛋白质结构倡议)和其他应用合作,如对序列与新确定的结构相似的蛋白质进行同源建模,为不完全结构建立完整的模型,以及预测潜在的突变位点以使蛋白质可溶。
英文摘要
DESCRIPTION (provided by applicant): The long-term objective of the proposed project is to provide a comprehensive platform, MUFOLD, for efficient and consistently accurate protein tertiary structure prediction. MUFOLD will help experimental biologists understand structures and functions of the proteins of their interest thereby facilitating hypotheses for experimental design. We will focus on the Funding Opportunity Announcement's second objective -- "High- Accuracy Models for Remote Homologs of Known Structures" which states "the quality of these models should be close to X-ray structures or high-resolution NMR structures with less than 2 Angstrom RMSD for backbone and side-chain atoms consistently for all protein targets." Specifically, we will integrate bioinformatics techniques, graph and network theories, computational algorithms, global optimization methods, statistics evaluations, etc. to develop a template-based structure prediction system, in which model generation, model quality assessment (QA), and model refinement will be seamlessly integrated together. At first, we will apply relevant information from the known template database (PDB) in depth as well as multi-layer QA methods to guide an efficient model generation in a small and targeted conformation space, which will facilitate computational efficiency and a limited number of models for QA methods to select. Secondly, we will improve the overall discerning power of QA by integrating various QA scores of a model and its structural relationships to other models generated for the same target protein. Thirdly, we will develop a population-based model refinement protocol, which integrates different levels of QA and efficient model generation techniques to improve the overall quality of models. Our goals are 1) to improve the prediction speed such that the prediction for a target protein with 200~300 residues can be finished in minutes on a multi-core desktop machine; 2) to enhance the QA ability of selecting the best models from the generated candidates, and decrease the current average ~10-point GDT-TS loss from the best available model to <5 points; 3) to achieve the prediction accuracy for remote homolog proteins within 2 Angstrom RMSD for backbone and side-chain atoms on average; and 4) to collaborate with PSI (Protein Structure Initiative) and others for various applications, such as performing homolog modeling for proteins with sequence similarity to newly determined structures, building complete models for incomplete structures, and predicting potential mutation sites to make protein soluble.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Multi-view self-supervised deep learning for biological sequences and beyond
  • 批准号:
    10623063
  • 项目类别:
  • 资助金额:
    $39.13万
  • 财政年份:
    2018
  • 负责人:
    DONG XU
  • 依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
  • 批准号:
    10395451
  • 项目类别:
  • 资助金额:
    $45.64万
  • 财政年份:
    2018
  • 负责人:
    DONG XU
  • 依托单位:
Interpretable and extendable deep learning model for biological sequence analysis and prediction
Deep learning for protein subcellular/sub-organelle localizations and localization motifs
海外基金