课题基金 / 基金详情

Collaborative Research: MFB: Integrating Deep Learning and High-throughput Experimentation to Rapidly Navigate Protein Fitness Landscapes for Non-native Enzyme Catalysis

Collaborative Research: MFB: Integrating Deep Learning and High-throughput Experimentation to Rapidly Navigate Protein Fitness Landscapes for Non-native Enzyme Catalysis
合作研究:MFB:整合深度学习和高通量实验,快速探索非天然酶催化的蛋白质适应性景观
批准号:
2226383
负责人:
Philip Romero
金额:
$48.99万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-11-01 至 2025-10-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
了解蛋白质结构和功能之间的关系仍然是一个重大挑战。这些知识将有利于药物设计、回收和化学生产。这个项目的目的是学习如何创造蛋白质,将促进在自然界中看到的反应。人工智能将解释实验产生的数据。两类酶将被修改以促进新的反应。为了使STEM劳动力多样化,将为对蛋白质设计感兴趣的学生提供机器学习研讨会。夏季研究机会将提供给传统上在STEM领域代表性不足的高中生和本科生。在这个项目中,蛋白质工程被视为一个贝叶斯优化问题,目的是探索序列空间以提高比活性。这种方法既模拟了预期的活动,也模拟了所作预测的不确定性。训练深度学习模型是数据密集型的。采用变压器结构的卷积神经网络(CNN)将使用模拟的序列函数数据进行预训练。模拟数据将使用Rosetta生成。预训练的CNN将使用组合密码子突变(CCM)生成的实验数据进行细化。将使用GFP表达、基于facs的筛选和下一代DNA测序来确定相应的氨基酸序列来监测单个细菌细胞中的酶活性。当存在多个细胞时,生物传感器筛选会受到串扰的影响。罗梅罗实验室开发的皮升级微滴筛选技术将用于避免这一问题。将开发一种模拟退火算法来随机搜索序列位置和退化密码子,以实现期望的批量BO目标。此外,还设计并实现了一个基于抽样推理的概率程序来估计密码子的最优组合。该项目由化学、生物工程、环境和运输系统部(CBET)、化学部(CHE)和信息和智能系统部(IIS)联合支持。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Understanding the relationship between protein structure and function remains a major challenge. This knowledge would benefit drug design, recycling, and chemical production. This project is designed to learn how to create proteins that will facilitate reactions seen in nature. Artificial intelligence will interpret the data generated by experiments. Two classes of enzymes will be modified to facilitate novel reactions. To help diversify the STEM workforce, workshops in machine learning will be offered to students interested in protein design. Summer research opportunities will be offered to high school and undergraduate students traditionally underrepresented in STEM fields.In this project, protein engineering is treated as a Bayesian optimization problem, with the objective to explore sequence space for improved specific activity. This approach models both the expected activity and the uncertainty of the prediction made. Training deep learning models is data intensive. A convolution neural net (CNN) using transformer architecture will use simulated sequence-function data to pretrain. The simulated data will be generated using Rosetta. Pretrained CNN will be refined with experimental data generated using combinatorial codon mutagenesis (CCM). Enzyme activity in single bacterial cells will be monitored using GFP expression, FACS-based screening, and next-generation DNA sequencing to determine the corresponding amino acid sequences. Biosensor screening can suffer from crosstalk when multiple cells are present. A picoliter-scale microdroplet screening technology developed in the Romero lab will be utilized to avoid this issue. A simulated annealing algorithm to randomly search over sequence positions and degenerate codons for libraries with high values for the expected batch BO objective will be developed. In addition, a probabilistic program using sampling-based inference to estimate the optimal combination of codons will be designed and implemented.This project is jointly supported by the Division of Chemical, Bioengineering, Environmental and Transport Systems (CBET), the Division of Chemistry (CHE), and the Division of Information and Intelligent Systems (IIS).This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Research on Quantum Field Theory without a Lagrangian Description
  • 批准号:
    24ZR1403900
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
    SATOSHI NAWATA
  • 依托单位:
Cell Research
Cell Research
Cell Research (细胞研究)