Machine learning approaches for faster discovery and adaptation of enzymes for difficult chemical reactions (MacBioSyn). Part I: providing solutions for regioselective oxygenations by 2OGD oxidases
Machine learning approaches for faster discovery and adaptation of enzymes for difficult chemical reactions (MacBioSyn). Part I: providing solutions for regioselective oxygenations by 2OGD oxidases
批准号:
497207454
负责人:
Dr. Mehdi Davari Dolatabadi, Ph.D.
金额:
$0.0万
依托单位国家:
德国
项目类别:
Priority Programmes
财政年份:
--
资助国家:
德国
项目状态:
未结题
起止时间:
中文摘要
生物催化合成化学物质被认为是未来绿色和可持续化学的基石。与欧盟委员会的数字化转型(绿色协议)相结合,这一点尤为突出。然而,它的力量在今天的工业中还远远没有实现,主要是因为可获得的酶的活性或多样性有限。2-氧戊二酸依赖(2OGD)蛋白是一个研究不足的酶家族,它催化“棘手的”氧化反应(例如,非活性炭的氧化官能化,去甲基化),这是具有挑战性的,或者不能用传统的化学合成来完成。因此,作为一种针对特定区域和产品的“化学替代品”,2OGD蛋白具有巨大的潜力,可以彻底改变该行业。确定这个大家族的新代表,例如增加底物范围,可以提供一系列新的生物催化途径,例如天然产物。然而,酶开发的一个共同挑战是通过基因组挖掘探索巨大的生物多样性来预测酶的活性。机器学习(ML)可以利用大量和多样化的酶数据集来预测功能和活性,并探索生物多样性以识别先进的生物催化剂。此外,机器学习方法可以同时优化多个蛋白质特性,并有效地导航序列和化学空间。在MacBioSyn项目中,我们的目标是开发一个通用的,高通量(HT)的基于ml的框架(深度学习,主动学习,强化学习)来预测酶的活性及其底物/反应范围。我们将通过协同方法,结合计算设计/建模(Davari)的跨学科专业知识和HT酶表征(Dippe/Wessjohann),基于筛选结果训练的ML模型,实现一个新的计算机框架,用于分析酶序列/底物对。为了建立这个平台,我们将专注于2OGD酶作为概念验证的生物催化剂。通过机器学习项目出现的关键挑战是可用于训练的数据集大小,以建立一致和可靠的统计建模。因此,MacBioSyn的目标是通过HT筛选(> 1500个酶)产生一个大的超家族生物多样性数据集。转换30个覆盖不同结构的基板将产生代表性数据,以迭代过程训练我们的算法。从本质上讲,我们的框架将为生物催化剂的发现提供活性/底物范围预测的解决方案。协同方法将提供方法,使ML方法的力量能够加速改进酶的发现,即如何开发生物催化反应(这里是氧化官能化)。从2OGD酶中学习到的新的基本设计原则将扩大其在有价值的天然产物的生物催化生产等方面的应用。
英文摘要
Biocatalytic synthesis of chemicals is considered a keystone for future green and sustainable chemistry. It is particularly highlighted in combination with digital transformation (Green Deal) by the European Commission. However, its power is far from being realized today in industry, mainly because of the limited activity or diversity of accessible enzymes. 2-oxoglutarate-dependent (2OGD) proteins are an under-researched family of enzymes which catalyze “tricky” oxidative reactions (e.g., oxyfunctionalization of non-activated carbons, demethylations), which are challenging or cannot be performed using traditional chemosynthesis. Thus, 2OGD proteins have the high potential to revolutionize the industry as a regio- and product-specific “alternative to chemistry”. Identifying new representatives of this large family having e.g. increased substrate scope can offer a new range of biocatalytic routes to e.g. natural products. However, a common challenge for enzyme development is the prediction of activity by exploring the enormous biodiversity through genome mining. Machine learning (ML) can capitalize on large and diverse enzyme datasets to predict function and activity, and explore the biodiversity to identify advanced biocatalysts. Additionally, ML methods comprise can optimize multiple protein traits simultaneously and to navigate sequence and chemical space efficiently.In the MacBioSyn project, we aim to develop (a general, high-throughput (HT)) ML-based framework (deep learning, active learning, reinforcement learning) that predicts the activity of enzymes and their substrate / reaction scope. We will implement a new in silico framework for the analysis of enzyme sequences/substrates pairs based on ML models trained on screening results by a synergistic approach, combining the interdisciplinary expertise of computational design / modeling (Davari) with HT enzyme characterization (Dippe/Wessjohann). To establish this platform, we will focus on 2OGD enzymes as proof-of-concept biocatalysts. The critical challenge that appears through machine learning projects is the dataset size available for training to establish consistent and reliable statistical modeling. Therefore, MacBioSyn aims at generating a large dataset by HT screening (> 1500 enzymes) of the superfamily’s biodiversity. Conversion of 30 substrates covering various structures will generate representative data to train our algorithms in an iterative process. In essence, our framework will provide a solution for activity / substrate scope prediction for biocatalyst discovery in general. The synergistic approach will provide methodologies that enable the power of ML methods to accelerate the discovery of improved enzymes, i. e. how biocatalytic reactions (here oxyfunctionalizations) are developed. The new fundamental design principles learned for 2OGD enzymes will broaden their applications in the biocatalytic production of valuable natural products and beyond.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Understanding the sequence-structure-function relationship of the large arylsulfate sulfotransferase (ASST) enzyme family for engineering novel sulfation biocatalysts
-
批准号:505682627
-
项目类别:Research Grants
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Dr. Mehdi Davari Dolatabadi, Ph.D.
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
儿童音乐能力发展对语言与社会认知能力及脑发育的影响
-
批准号:31971003
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:南云
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
多场景网络学习中基于行为-情感-主题联合建模的学习者兴趣挖掘关键技术研究
-
批准号:61702207
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2017
-
负责人:刘智
-
依托单位:
基于异构医学影像数据的深度挖掘技术及中枢神经系统重大疾病的精准预测
-
批准号:61672236
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2016
-
负责人:王骏
-
依托单位: