Machine learning approaches for the discovery, repurposing, and optimization of natural products with therapeutic potential
Machine learning approaches for the discovery, repurposing, and optimization of natural products with therapeutic potential
批准号:
10693375
负责人:
Allison Sara Walker
金额:
$39.63万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-09-01 至 2027-08-31
关键词:
AnabolismAttentionBacteriaBehaviorBiochemicalBiologicalChemical StructureChemicalsComplexComputational TechniqueComputersCouplingDataDirected Molecular EvolutionEnzymesGene ClusterGenesGenetic StructuresMachine LearningMethodsModelingMolecular MachinesNatural ProductsPathway interactionsPeptide Leader SequencesPeptidesPlantsPropertyRibosomesSourceStructureStructure-Activity RelationshipSystemTherapeuticTherapeutic UsesTrainingWorkanalogcomputerized toolsdesigndrug discoveryexperimental studyfungusgene productgenetic approachgraph neural networkhigh throughput screeningimprovedinhibitorinterestmachine learning algorithmmachine learning methodmachine learning modelmethod developmentmolecular modelingnatural product inspirednovelpeptide synthasepolyketide synthaseprotein aminoacid sequenceprotein protein interactiontrend
中文摘要
项目摘要
长期以来,细菌、真菌和植物的天然产物一直是有用分子的丰富来源。然而,
由于其复杂的结构,很难筛选出许多天然产物的类似物来真正了解
管理其结构和活动之间关系的规则。我们将通过发展
一种能够对天然化合物的结构-活性关系(SAR)进行功能建模的机器学习方法
并协助设计可合成天然产物类似物的生物合成途径。因此,
我们将开发两种方法来优先考虑最有可能作为治疗药物的天然产品
用于活性筛选和生物合成感兴趣的天然产物。
机器学习是一种强大的计算技术,使计算机能够从数据中进行推理。
生物分子有丰富的序列、结构和活性数据,我们可以用它们来
建立机器学习模型来预测生化系统的行为。平整机
学习不完全准确的算法对药物发现工作可能非常有用。这是有可能的
使用机器学习筛选比高通量筛选多几个数量级的化合物。
因此,机器学习可以用作提高屏幕命中率的初始过滤器。
我们的第一个项目将应用机器学习来研究天然产物SARS。我们将采取两种方法,一是
遗传和化学结构方法。在遗传学方法中,我们将验证
我们之前观察到的生物合成基因和天然产物活性之间的相关性
延伸到由生物合成基因安装的化学亚结构。在化学结构方法中,我们
将研究图形神经网络预测天然产品性质的能力。
我们的第二个项目将专注于开发机器学习和其他用于设计的计算工具
生物合成基因簇(BGC)用于生物合成新的天然产物类分子。我们首先要关注的是
核糖体合成和翻译后修饰多肽(RIPP)及其预测方法的发展
相容的修饰酶-前导肽对。为了做到这一点,我们将使用分子建模、统计
耦合分析(SCA)和机器学习。在RIPPS上验证了我们的方法之后,我们将把注意力转向
对于更难的BGC类,如非核糖体肽合成酶(NRPS)和聚酮合成酶
(PKS)。
我们的第三个项目是开发设计基于RIPP的蛋白质-蛋白质相互作用(Ppi)的方法。
抑制剂。我们将开发分子建模和机器学习方法来预测最佳RIPP
抑制感兴趣的PPI的序列。然后,我们将验证并收集以下项目的其他培训数据
使用定向进化实验进行预测。
英文摘要
Project Summary
Natural products from bacteria, fungi, and plants have long been a rich source of useful molecules. However,
due to their complex structures, it is difficult to screen many analogs of natural products to truly understand the
rules governing the relationship between their structure and activity. We will address this challenge by developing
machine learning methods that can functionally model the structure-activity relationships (SAR) of natural
products and aid in the design of biosynthetic pathways that can synthesize natural product analogs. Therefore,
we will develop methods both for prioritizing natural products that are most likely to be useful as therapeutics for
activity screens and for biosynthesizing natural products of interest.
Machine learning is a powerful computational technique that enables computers to make inferences from data.
There is a wealth of sequence, structure, and activity data available for biological molecules that we can use to
build machine learning models to make predictions about the behavior of biochemical systems. Even machine
learning algorithms that are not perfectly accurate can be extremely useful for drug discovery efforts. It is possible
to screen orders of magnitude more compounds using machine learning than in high-throughput screens.
Machine learning can therefore be used as an initial filter to increase hit rates in screens.
Our first project will apply machine learning to study natural product SARs. We will take two approaches, a
genetic and chemical structure approach. In the genetic approach we will validate correlations between
biosynthetic genes and natural product activity that we have previously observed and confirm that the correlation
extends to chemical substructures installed by the biosynthetic genes. In the chemical structure approach, we
will investigate the ability of graph neural networks to predict natural product properties.
Our second project will focus on developing machine learning and other computational tools for designing
biosynthetic gene clusters (BGCs) to biosynthesize novel natural product-like molecules. We will first focus on
Ribosomally Synthesized and Posttranslationally modified Peptides (RiPPs) and develop methods to predict
compatible modifying enzyme-leader peptide pairs. To do this we will use molecular modeling, Statistical
Coupling Analysis (SCA), and machine learning. After validating our methods on RiPPs, we will turn our attention
to more difficult classes of BGCs, such as nonribosomal peptide synthetases (NRPS) and polyketide synthases
(PKS).
Our third project is the development of methods for designing RiPP-based protein-protein interaction (PPI)
inhibitors. We will develop both molecular modeling and machine learning methods for predicting optimal RiPP
sequences for inhibiting a PPI of interest. We will then validate and collect additional training data for these
predictions using directed evolution experiments.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Equipment Supplement to R35GM146987: Purchase of LC-MS system for high throughput isolation of bioactive natural products
-
批准号:10798569
-
项目类别:
-
资助金额:$24.19万
-
财政年份:2022
-
负责人:Allison Sara Walker
-
依托单位:
Bioinformatics and Chemical Biology Approaches for Identifying Bioactive Natural Products of Symbiotic Actinobacteria
-
批准号:9540546
-
项目类别:
-
资助金额:$5.83万
-
财政年份:2018
-
负责人:Allison Sara Walker
-
依托单位:
国内基金
海外基金
多模态超声VisTran-Attention网络评估早期子宫颈癌保留生育功能手术可行性
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:郑巧
-
依托单位:
Ultrasomics-Attention孪生网络早期精准评估肝内胆管癌免疫治疗的研究
-
批准号:--
-
项目类别:面上项目
-
资助金额:52万元
-
批准年份:2022
-
负责人:陈立达
-
依托单位: