Machine learning of biomolecular interactions and the human signaling networks they comprise
Machine learning of biomolecular interactions and the human signaling networks they comprise
批准号:
10714785
负责人:
Mohammed Nazar AlQuraishi
金额:
$41.13万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-22 至 2028-08-31
关键词:
AddressAffinityAllelesAmino Acid SequenceBiologicalCellsChemicalsCollaborationsComputer ModelsDiseaseGenetic DiseasesGenomeHumanHuman BiologyIndividualInformation NetworksLanguageLearningLigand BindingLigandsMachine LearningMapsModelingMutationNational Heart, Lung, and Blood InstitutePathogenicityPathway interactionsPersonsPopulationPropertyProtein ConformationProteinsProteomeResearchSignal TransductionStructureSystems BiologyTechniquesTestingTrans-Omics for Precision MedicineVariantbiobankcomputerized toolsdeep learningexomegenetic analysisgenomic datahuman diseaseinformation processinglearning strategymachine learning modelmathematical learningpredictive modelingprogramsprotein protein interactionprotein structure predictionsynergismtrait
中文摘要
我的实验室将使用机器学习来建立生物分子及其相互作用的物理基础模型
在蛋白质组(基因组)尺度上应用这些模型来解决人类系统生物学中的基本问题
发信号。在建模方面,我们将致力于建立蛋白质-配体的计算模型
相互作用,特别强调翻译后修饰的配体,细胞广泛用于信号传递
网络。我假设,蛋白质-配体相互作用模型的准确性和通用性的一步变化是
可能使用深度学习在蛋白质结构预测和蛋白质表示学习方面的进展。我的
Lab一直走在这些进步的前沿,开发了第一个端到端可区分的模型
蛋白质结构预测(RGN);第一个蛋白质语言模型(UniRep),学习的关键技术
捕捉蛋白质的化学、结构和进化属性的数学表示;以及一个
最早的蛋白质-蛋白质相互作用(HSM)深度学习方法之一。我们将利用我们在这些方面的专业知识
基于序列和结构信息预测蛋白质-配体相互作用的结构域。我们将进一步
为此,开发专门的模型来预测蛋白质结构和变化的蛋白质构象
预测蛋白质-配体相互作用,使用这些预测作为我们的蛋白质-配体相互作用模型的输入。
在生物学方面,我们将使用这些机器学习的模型来组装特定于人的信号
了解正常的等位基因变异是如何在信令网络水平上表现出来的,以及如何
这些网络在人类疾病中受到干扰。为了研究信令网络中的一般变化,我们将使用
外显子组序列(UK Biobank和NHLBI TOPMed),以构建映射特定人的个性化网络
蛋白质序列与蛋白质配体亲和力的关系。我们将量化不同个体之间的网络拓扑差异
和种群,并测试疾病相关特征是否与拓扑相关。我们还将比较
健康人和疾病患者的网络,以确定易患个人的拓扑差异
遗传性疾病。最终,我希望机器学习的模型能够充分预测配基结合
通过突变对通路重新连接的机械论理解是可能的。虽然我的重点将是
计算,我希望进行密切合作-与Fordyce实验室(斯坦福大学)进行实验
表征和验证蛋白质-配体相互作用和沈实验室(哥伦比亚)执行统计遗传
分析-在计算和实验的界面上开发协同效应。
英文摘要
My lab will use machine learning to build physically-grounded models of biomolecules and their interactions and
apply these models at proteome (genome) scale to address basic questions in the systems biology of human
signaling. On the modeling front, our efforts will focus on building computational models of protein-ligand
interactions, with a specific emphasis on post-translationally modified ligands that cells widely employ in signaling
networks. I hypothesize that a step change in accuracy and generality of protein-ligand interaction models is
possible using deep learning advances in protein structure prediction and protein representation learning. My
lab has been at the forefront of these advances, having developed the first end-to-end differentiable model of
protein structure prediction (RGN); the first protein language model (UniRep), a key technique for learning
mathematical representations that capture chemical, structural, and evolutionary properties of proteins; and one
of the first deep learning methods for protein-protein interactions (HSM). We will leverage our expertise in these
domains to predict protein-ligand interactions based on both sequence and structure information. We will further
develop specialized models for predicting protein structures and alternate protein conformations for the purpose
of predicting protein-ligand interaction, using these predictions as inputs for our protein-ligand interaction models.
On the biological front, we will employ these machine-learned models to assemble person-specific signaling
networks to understand how normal allelic variation is manifested at the level of signaling networks, and how
these networks are perturbed in human diseases. To study general variation in signaling networks, we will use
exome sequences (UK Biobank and NHLBI TOPMed) to build individualized networks that map person-specific
protein sequences to protein-ligand affinities. We will quantify how network topology varies among individuals
and populations and test whether disease-associated traits correlate with topology. We will also compare
networks of healthy and disease-afflicted persons to identify topological differences that predispose individuals
to genetic diseases. Ultimately, I expect machine-learned models to be sufficiently predictive of ligand binding
that mechanistic understanding of pathway rewiring by mutations is possible. While my focus will be
computational, I expect to carry out close collaborations—with the Fordyce Lab (Stanford) to experimentally
characterize and validate protein-ligand interactions and the Shen Lab (Columbia) to perform statistical genetic
analyses—to exploit synergies at the interface of computation and experimentation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金