课题基金 / 基金详情

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
我的实验室将使用机器学习来建立生物分子及其相互作用的物理基础模型, 在蛋白质组(基因组)规模上应用这些模型来解决人类系统生物学中的基本问题, 信号在建模方面,我们的工作将集中在建立蛋白质-配体的计算模型 相互作用,特别强调细胞广泛用于信号传导的后修饰配体 网络.我假设,蛋白质-配体相互作用模型的准确性和通用性的阶跃变化是 可能使用深度学习在蛋白质结构预测和蛋白质表示学习方面的进展。我 实验室一直处于这些进步的最前沿,已经开发了第一个端到端的可区分模型, 蛋白质结构预测(RGN);第一个蛋白质语言模型(UniRep),学习的关键技术 捕获蛋白质的化学、结构和进化特性的数学表示;以及 蛋白质相互作用(HSM)的第一个深度学习方法。我们将利用我们的专业知识, 结构域来预测基于序列和结构信息的蛋白质-配体相互作用。我们将进一步 开发用于预测蛋白质结构和替代蛋白质构象的专用模型, 预测蛋白质-配体相互作用,使用这些预测作为我们的蛋白质-配体相互作用模型的输入。 在生物学方面,我们将使用这些机器学习模型来组装个人特定的信号 网络,以了解正常的等位基因变异是如何表现在信号网络的水平,以及如何 这些网络在人类疾病中受到干扰。为了研究信令网络的一般变化,我们将使用 外显子组序列(英国生物库和NHLBI TOPMed),以建立个性化的网络,映射个人特异性 蛋白质序列与蛋白质-配体亲和力的关系。我们将量化网络拓扑在个体之间的变化 和群体,并测试疾病相关性状是否与拓扑结构相关。我们还将比较 健康人和患病者的网络,以确定易使个人患病的拓扑差异 遗传疾病。最终,我希望机器学习模型能够充分预测配体结合 通过突变对通路重新布线的机械理解是可能的。而我的重点将是 计算,我希望进行密切合作,与福代斯实验室(斯坦福大学),以实验 表征和验证蛋白质-配体相互作用和Shen Lab(哥伦比亚)进行统计遗传学分析 分析-利用计算和实验界面的协同作用。
英文摘要
My lab will use machine learning to build physically-grounded models of biomolecules and their interactions and apply these models at proteome (genome) scale to address basic questions in the systems biology of human signaling. On the modeling front, our efforts will focus on building computational models of protein-ligand interactions, with a specific emphasis on post-translationally modified ligands that cells widely employ in signaling networks. I hypothesize that a step change in accuracy and generality of protein-ligand interaction models is possible using deep learning advances in protein structure prediction and protein representation learning. My lab has been at the forefront of these advances, having developed the first end-to-end differentiable model of protein structure prediction (RGN); the first protein language model (UniRep), a key technique for learning mathematical representations that capture chemical, structural, and evolutionary properties of proteins; and one of the first deep learning methods for protein-protein interactions (HSM). We will leverage our expertise in these domains to predict protein-ligand interactions based on both sequence and structure information. We will further develop specialized models for predicting protein structures and alternate protein conformations for the purpose of predicting protein-ligand interaction, using these predictions as inputs for our protein-ligand interaction models. On the biological front, we will employ these machine-learned models to assemble person-specific signaling networks to understand how normal allelic variation is manifested at the level of signaling networks, and how these networks are perturbed in human diseases. To study general variation in signaling networks, we will use exome sequences (UK Biobank and NHLBI TOPMed) to build individualized networks that map person-specific protein sequences to protein-ligand affinities. We will quantify how network topology varies among individuals and populations and test whether disease-associated traits correlate with topology. We will also compare networks of healthy and disease-afflicted persons to identify topological differences that predispose individuals to genetic diseases. Ultimately, I expect machine-learned models to be sufficiently predictive of ligand binding that mechanistic understanding of pathway rewiring by mutations is possible. While my focus will be computational, I expect to carry out close collaborations—with the Fordyce Lab (Stanford) to experimentally characterize and validate protein-ligand interactions and the Shen Lab (Columbia) to perform statistical genetic analyses—to exploit synergies at the interface of computation and experimentation.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金