课题基金 / 基金详情

MFB: Deep-Learning Enabled Structure Prediction and Design of Protein-DNA Assemblies

MFB: Deep-Learning Enabled Structure Prediction and Design of Protein-DNA Assemblies
MFB:深度学习支持蛋白质-DNA 组装的结构预测和设计
批准号:
2226466
负责人:
David Baker
金额:
$149.85万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-01 至 2025-08-31

项目摘要

项目成果

David Baker的其他基金

相似基金

相关文献

中文摘要
翻译
在这个生物技术分子基础(MFB)项目中,华盛顿大学生物化学系的David Baker和Frank DiMaio教授,以及Fred Hutchinson癌症中心基础科学的Barry Stoddard教授正在共同开发利用深度学习(DL)方法建模和设计蛋白质- dna复合物的新方法。为此,他们将开发三个基于dl的模型:(1)从序列中预测蛋白质- dna复合体结构的模型,(2)蛋白质- dna复合体序列设计的模型,以及(3)蛋白质- dna复合体结构预测和设计的质量评估模型。本项目开发的DL模型将用于设计序列特异性DNA结合微型蛋白,能够靶向特定的dsDNA序列。在这种情况下,使用新的基于dl的模型将有助于验证模型的准确性,并将作为设计生物技术应用的蛋白质- dna界面的强大工具产生广泛的影响,例如设计新的转录因子,核酸修饰酶和基因校正试剂。该项目是DL研究、计算蛋白设计、生物化学和结构生物学的交叉点,将为参与该项目的本科生、研究生和博士后提供多学科培训。他们的推广和教育项目的主要目标是吸引年轻人从事STEM(科学、技术、工程和数学)领域的职业,并改善生物化学和计算蛋白质设计方面的培训。外展计划包括多管齐下的努力,重点是通过个人指导的夏季研究和一个将在学年期间运行的基于队列的本科生研究项目来吸引本科生。这两方面的努力都将集中在培养本科生学习计算蛋白质设计的当代方法和验证蛋白质功能的实验方法,包括本提案中开发的新方法。该项目旨在开发一套机器学习/深度学习(ML/DL)技术来模拟蛋白质- dna复合物。能够推断蛋白质- dna复合物结构、预测dna结合蛋白(DBPs)的核苷酸特异性以及评估蛋白质- dna复合物模型准确性的新工具对于解决突出的技术问题(如开发新的转录因子)将是无价的。目前的方法缺乏准确性或计算量大,主要是由于在模拟DNA构象灵活性、氢键和静电相互作用、金属离子辅助因子和蛋白质-DNA复合物的高度溶剂化界面的间接读出方面存在困难。具体目标是开发基于ML的方法,用于(1)基于最近开发的用于预测蛋白质结构的机器学习框架RoseTTAFold模型,从序列和序列比对中推断DNA和蛋白质-DNA复合物的结构模型;(2)基于蛋白质- dna复合物主链信息设计序列特异性dbp并预测其特异性的序列预测神经网络;(3)评价蛋白质- dna复合物结构模型的精度预测模型。本项目开发的三种深度学习方法将在dbp的设计中得到利用。设计的DBPs将在高通量池格式下进行实验验证,使用酵母展示、细胞分选和下一代测序方法来近似池设计的结合亲和力。在酵母展示实验中显示DNA结合活性的设计将使用体外生物化学技术进一步表征DNA结合亲和力和特异性,设计模型将通过x射线共结晶验证。在这种设计背景下,ML模型的应用将验证模型的准确性,并为设计生物技术应用的蛋白质- dna界面提供一个强大的工具,如设计新的转录因子、核酸修饰酶和基因校正试剂。本项目由化学部(CHE)、信息与智能系统部(IIS)和物理部(PHY)联合支持。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
In this Molecular Foundation for Biotechnology (MFB) project, Professors David Baker and Frank DiMaio of the Department of Biochemistry at the University of Washington, and Barry Stoddard of Basic Sciences at the Fred Hutchinson Cancer Center together are developing new ways to model and design protein-DNA complexes using deep-learning (DL) methods. To do this, they will develop three DL-based models: (1) a model for prediction of protein-DNA complex structures from sequence, (2) a model for sequence design of protein-DNA complexes, and (3) a model for quality assessment of protein-DNA complex structure predictions and designs. The DL models developed in this project will be leveraged in a pipeline for design of sequence-specific DNA binding miniproteins capable of targeting specified sequences of dsDNA. The use of the novel DL-based models in this context will be useful for validating model accuracy and will have broad impact as a powerful tool for designing protein-DNA interfaces for biotechnology applications, such as the design of novel transcription factors, nucleic acid modifying enzymes, and gene correction reagents. This project lies at the interface of DL research, computational protein design, biochemistry, and structural biology and will provide multi-disciplinary training for undergraduates, graduate students, and postdocs involved in the project. The primary goals of their outreach and education programs are to attract young people to careers in STEM (science, technology, engineering and mathematics) and improve training in biochemistry and computational protein design. The outreach plan involves a multi-pronged effort focused on engaging undergraduates through individually mentored summer research and a cohort-based undergraduate research program that will run during the academic year. Both efforts will be focused on training undergraduates in contemporary methods in computational protein design and experimental methods for validating protein function, including the novel methods developed in this proposal. This project seeks to develop a suite of machine learning/deep learning (ML/DL) techniques for modeling protein-DNA complexes. New tools capable of inferring protein-DNA complex structures, predicting the nucleotide specificity of DNA-binding proteins (DBPs), and evaluating accuracy of protein-DNA complex models would be invaluable in solving salient technological problems, such as developing novel transcription factors. Current approaches lack accuracy or are computationally intensive, primarily due to the difficulties in modeling indirect readout of DNA conformational flexibility, hydrogen bonding and electrostatic interactions, metal ion cofactors, and the highly solvated interfaces of protein-DNA complexes. The specific goals are to develop DL-based methods for (1) Inference of structure models of DNA and protein-DNA complexes from sequences and sequence alignments, based on the recently developed RoseTTAFold model, an ML framework for predicting protein structures; (2) A sequence prediction neural network for designing sequence specific DBPs and predicting their specificity given protein-DNA complex backbone information, and (3) An accuracy prediction model for evaluating structural models of protein-DNA complexes. The three DL methods developed in this project will be leveraged in the design of DBPs. Designed DBPs will be experimentally validated in a high-throughput pooled format using yeast display, cell sorting, and next-generation sequencing methods to approximate the binding affinity of pooled designs. Designs showing DNA binding activity in yeast display experiments will be further characterized for DNA binding affinity and specificity using in vitro biochemistry techniques and the design models will be confirmed with X-ray co-crystallization. Application of the ML models in this design context will provide validation of model accuracy and result in a powerful tool for designing protein-DNA interfaces for biotechnology applications, such as the design of novel transcription factors, nucleic acid modifying enzymes, and gene correction reagents. This project is jointly supported by the Division of Chemistry (CHE), the Division of Information and Intelligent Systems (IIS), and the Division of Physics (PHY).This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Co-production of a software tool for field-scale species distribution modelling (fs-SDM) and mapping using local biodiversity records
  • 批准号:
    NE/V007726/1
  • 项目类别:
    Fellowship
  • 资助金额:
    $25.61万
  • 财政年份:
    2020
  • 负责人:
    David Baker
  • 依托单位:
CIBR: Collaborative Research: CIBR Expanding structure coverage of genomes to facilitate macromolecular assembly determination.
  • 批准号:
    1937533
  • 项目类别:
    Standard Grant
  • 资助金额:
    $47.36万
  • 财政年份:
    2019
  • 负责人:
    David Baker
  • 依托单位:
Generation, functionalization, and distribution of de novo designed protein nanomaterials
  • 批准号:
    1629214
  • 项目类别:
    Standard Grant
  • 资助金额:
    $135.0万
  • 财政年份:
    2016
  • 负责人:
    David Baker
  • 依托单位:
RAPID: Empowering the Citizen Scientist in the Fight Against Ebolaviruses
  • 批准号:
    1523362
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2015
  • 负责人:
    David Baker
  • 依托单位:
国内基金
海外基金
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
  • 批准号:
    2026JJ81909
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    胡曦
  • 依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
  • 批准号:
    12271434
  • 项目类别:
    面上项目
  • 资助金额:
    46万元
  • 批准年份:
    2022
  • 负责人:
    贺小伟
  • 依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
  • 批准号:
    2020A151501709
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2020
  • 负责人:
    谢怡
  • 依托单位:
面向Deep Web的数据整合关键技术研究
  • 批准号:
    61872168
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2018
  • 负责人:
    董永权
  • 依托单位: