课题基金 / 基金详情

Next-Generation Protein Engineering: Machine Learning for Enzyme Engineering

Next-Generation Protein Engineering: Machine Learning for Enzyme Engineering
下一代蛋白质工程:酶工程的机器学习
批准号:
1937902
负责人:
Frances Arnold
金额:
$78.75万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-10-01 至 2024-09-30

项目摘要

项目成果

Frances Arnold的其他基金

相似基金

相关文献

中文摘要
翻译
酶是催化反应的蛋白质。它们被用于各种应用,如清洁污渍的洗涤剂、检测血糖水平的传感器、高果糖玉米糖浆生产等工业过程。酶可以通过一种叫做定向进化(DE)的方法来改进它们的性能。这是一个重复基因序列突变和酶筛选的过程,以确定突变序列引起的性能变化。DE实验的一个副产品是大量未使用的数据。该项目将把机器学习(ML)整合到DE工作流程中。这些数据和其他数据将训练机器学习算法来识别并最终预测蛋白质结构,这些结构对创建所需的酶活性有用。目标是能够以更快的速度和更低的成本制造新的有用的酶。生命的多样性在很大程度上源于蛋白质的进化和适应能力。20种蛋白质原氨基酸的排列允许超天文数字可能的蛋白质序列,其中绝大多数不折叠或编码有用的功能。我们提出机器学习(ML)模型可以使用从定向进化(DE)实验中获得的信息来提高在这个空间中搜索功能蛋白的效率。虽然诸如DE之类的技术隐式地依赖于蛋白质序列空间功能景观的底层结构,但显式地对该结构进行建模将允许更有效的搜索算法。我们最近展示了一种数据驱动的机器学习方法来指导DE实验,该方法解释了蛋白质突变的上位性,并使多个有益突变能够被纳入单代突变和筛选中。我们的目标是进一步发展这一工作流程,以同时解决多个任务,特别是预测跨多种底物的酶活性。为了实现这一目标,我们建议将有关多种底物的信息纳入ml引导的定向进化中,并使用为单独预测任务开发的蛋白质和底物的有效编码。然而,这些编码并没有优化到可以一起用于ML。因此,我们建议开发新的编码,以内聚和协同的方式描述酶系统的组成部分。最后,虽然定向进化已经成功地将酶应用于人类,但这一过程目前需要专家知识和每项工程任务的密集试错。当面临寻找DE起始活性的艰巨挑战时,对专家知识的要求是最明显的。使用ML,我们将建立一个模型,该模型可以基于已知的这些酶的天然底物来预测p450对目标底物的非天然碳/亚硝基烯转移活性。该奖项由化学、生物工程、环境和运输系统部门的细胞和生物化学工程项目以及分子和细胞生物科学部门的系统和合成生物学项目共同创立。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Enzymes are proteins that catalyze reactions. They are used in various applications such as detergents to clean stains, sensors to detect blood sugar levels, industrial processes such as production of high fructose corn syrup. Enzymes can be engineered to improve their performance using a method called directed evolution (DE). This is a process of repeated gene sequence mutation and enzyme screening to determine the change in performance caused by the mutated sequences. A side-product of DE experiments is an abundance of unused data. This project will incorporate machine learning (ML) into the DE workflow. This and other data will train the ML algorithm to recognize and ultimately predict protein structures that are useful in creating the enzyme activity desired. The objective is to be able to make novel and useful enzymes more rapidly and at lower cost.The diversity of life arises in great measure from the ability of proteins to evolve and adapt. The permutations of the 20 proteinogenic amino acids allow for supra-astronomical numbers of possible protein sequences, the vast majority of which do not fold or encode a useful function. We propose that machine learning (ML) models can use information gained from directed evolution (DE) experiments to improve the efficiency of searching this space for functional proteins. While techniques such as DE implicitly rely on an underlying structure of the functional landscape of protein sequence space, explicitly modeling this structure would allow for far more efficient search algorithms. We recently demonstrated a data-driven, ML approach to guiding DE experiments which accounts for the epistatic nature of protein mutations and enables multiple beneficial mutations to be incorporated in a single generation of mutation and screening. We aim to further develop this workflow to address multiple tasks simultaneously, specifically to predict enzyme activity across multiple substrates. To accomplish this, we propose to incorporate information about multiple substrates into ML-guided directed evolution and use validated encodings for proteins and substrates developed for separate predictive tasks. However, these encodings are not optimized to work together for ML. We therefore propose to develop new encodings that describe the components of an enzymatic system in a cohesive and synergistic manner. Finally, while directed evolution has successfully adapted enzymes for human applications, this process currently requires expert knowledge and intensive trial-and-error for each engineering task. The requirement for expert knowledge is clearest when approaching the formidable challenge of finding starting activity for DE. Using ML, we will build a model that can predict the non-natural carbene/nitrene transfer activities of P450s against target substrates, based on what is known of the natural substrate(s) of these enzymes.This award is cofounded by the Cellular and Biochemical Engineering Program in the Division of Chemical, Bioengineering, Environmental and Transport Systems and the Systems and Synthetic Biology Program in the Division of Molecular and Cellular Biosciences.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(5)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1021/acssynbio.1c00592
发表时间: 2022-03-18
期刊: ACS SYNTHETIC BIOLOGY
影响因子: 4.7
作者: [Wittmann, Bruce J., Johnston, Kadina E., Arnold, Frances H.]
通讯作者: Arnold, Frances H.
DOI: 10.1016/j.sbi.2021.01.008
发表时间: 2021-08-01
期刊: CURRENT OPINION IN STRUCTURAL BIOLOGY
影响因子: 6.8
作者: [Wittmann, Bruce J., Johnston, Kadina E., Arnold, Frances H.]
通讯作者: Arnold, Frances H.
Evolving Hemoproteins for New-to-Nature Ring-Forming Reactions
  • 批准号:
    2016137
  • 项目类别:
    Standard Grant
  • 资助金额:
    $120.14万
  • 财政年份:
    2020
  • 负责人:
    Frances Arnold
  • 依托单位:
Expanding the Enzyme Repertoire by Evolution and Engineering
  • 批准号:
    1513007
  • 项目类别:
    Standard Grant
  • 资助金额:
    $89.12万
  • 财政年份:
    2015
  • 负责人:
    Frances Arnold
  • 依托单位:
SusChEM: Engineering and Evolution of Cytochrome P450 Enzymes for Non-Natural Chemistry
  • 批准号:
    1403077
  • 项目类别:
    Standard Grant
  • 资助金额:
    $34.5万
  • 财政年份:
    2014
  • 负责人:
    Frances Arnold
  • 依托单位:
Collaborative Research: Metabolically Engineered Organisms for Conversion of Cellulose to Isobutanol
  • 批准号:
    0903817
  • 项目类别:
    Standard Grant
  • 资助金额:
    $54.3万
  • 财政年份:
    2009
  • 负责人:
    Frances Arnold
  • 依托单位:
国内基金
海外基金
Next Generation Majorana Nanowire Hybrids