Discovering interpretable mechanisms explaining high dimensional biomolecular data
Discovering interpretable mechanisms explaining high dimensional biomolecular data
批准号:
10711988
负责人:
Milo Lin
金额:
$41.0万
依托单位国家:
美国
项目类别:
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-09-01 至 2028-07-31
关键词:
AccelerationAddressAmino Acid SequenceAmyloid beta-ProteinAntibiotic ResistanceAntibioticsAutomobile DrivingBehaviorBiological AssayBiologyCell Fate ControlCollaborationsComplexComputing MethodologiesDataData SetDimensionsDirected Molecular EvolutionDiseaseFoundationsFutureGoalsHealthHumanInformaticsIntuitionLearningLibrariesLiquid substanceMethodsModelingNetwork-basedNeurobiologyNeurodegenerative DisordersPatientsPatternPeptide LibraryPeptidesPharmaceutical PreparationsPhysicsProteinsRNARNA SequencesResourcesRunningSamplingScienceStructureTimeTrainingWorkartificial neural networkbeta-Lactamasedeep learningdisease diagnosisexperimental studyhigh dimensionalityinsightinventionlarge datasetsmolecular dynamicsmonomermutantneural networkneurotoxicprotein aggregationself assemblysimulationtau Proteinstau aggregation
中文摘要
发现解释高维生物分子数据的可解释机制
项目总结。蛋白质和RNA序列如何编码折叠、聚集和功能是一个基本的问题
对人类健康有广泛影响的问题。发现该编码的预测原则
需要提供机械洞察力的计算方法,特别是对于大部分本质上
实验结构信息有限的无序蛋白质。然而,复杂性和维度
这一问题对现有的计算方法提出了根本性的挑战。公理的方法,
从第一性原理出发的建模行为受到仿真运行时间和未知上下文相关的限制
参数。深度学习等基于信息学的方法可能通过以下方式发现原理
跨规模和复杂性集成大型数据集。然而,这些模型产生了“黑箱”预测。
一)难以理解,二)在其训练数据之外概括得很差(即,很好地理解的制度)。
我的实验室开发了一些方法来克服这两种方法的局限性。(1)不言自明:我们
开发了一种统计物理方法,以指数级提高蛋白质自组装的采样
分子动力学模拟中的结构非均相单体。(2)信息学:我们创造了本质
基于神经生物学原理的神经网络(Enns),并证明了它们克服了上述问题
深度学习在广泛的学习任务上的局限性,包括序列到函数的预测。
在接下来的五年里,我的实验室将同时使用公理和信息学方法来处理三个实例
序列-结构-功能问题:1)使用增强的采样分子动力学模拟来
发现神经毒性低聚物的过渡态和Abeta和tau肽单体的纤维形成;2)用途
Enns将发现推动神经退行性变中RNA相关tau原纤维聚集的RNA序列规则
疾病利用tau蛋白和共定位的RNA序列数据集;3)使用Enns来提取序列规则
确定一株或突变的β-内酰胺酶蛋白是否能中和不同种类的
药物小组,并确定未来潜在的抗生素耐药突变株。我们的长期目标是开发一个新的网络--
基于平台的数据到公理的自动转换。利用成熟的协作
与具有广泛专业知识的同事一起,我们将通过结合我们独特的计算方法来实现这些目标
有了包括时间分辨蛋白质聚集分析在内的实验资源,患者来源的tau纤维共同-
定位于序列特定的RNA,高通量液体培养抗生素筛选,多重定向
抗生素耐药性的进化实验,以及大型内部多肽和RNA突变文库。
这项工作为将大数据集转换为人类可理解的规则奠定了基础
将序列与功能联系起来,并将这些规则与结构动力学的物理机制联系起来。此入站
TURN可以加速疾病的诊断和治疗。
英文摘要
Discovering interpretable mechanisms explaining high-dimensional biomolecular data
Project summary. How protein and RNA sequence encodes folding, aggregation, and function is a fundamental
question with wide-ranging human health implications. Discovering predictive principles for this encoding
requires computational approaches that offer mechanistic insight, especially for the large fraction of intrinsically
disordered proteins for which experimental structural information is limited. Yet the complexity and dimensionality
of this problem poses fundamental challenges to existing computational methods. The axiomatic approach,
modeling behavior from first-principles, is limited by simulation runtime and unknown context-dependent
parameters. Informatics-based approaches such as deep learning could potentially discover principles by
integrating large datasets across scales and complexity. However, these models produce “black box” predictions
that i) are difficult to understand and ii) generalize poorly beyond their training data (i.e. well-understood regime).
My lab developed methods to overcome limitations of both types of approaches. (1) Axiomatic: we
developed a statistical physics method to exponentially enhance sampling of protein self-assembly from
structurally heterogeneous monomers in molecular dynamics simulations. (2) Informatic: we invented essence
neural networks (ENNs) based on neurobiological principles and demonstrated that they overcome the above
limitations of deep learning on a wide range of learning tasks, including sequence-to-function prediction.
Using both axiomatic and informatic approaches, in the next five years my lab will tackle three instances
of the sequence-structure-function problem: 1) Use enhanced sampling molecular dynamics simulations to
discover transition states of neurotoxic oligomer and fibril formation of Abeta and tau peptide monomers; 2) Use
ENNs to discover the RNA-sequence rules driving RNA-associated tau fibril aggregation in neurodegenerative
disease using tau protein and colocalized RNA sequence datasets; 3) Use ENNs to distill the sequence rules
determining whether a strain or mutant of beta lactamase protein can neutralize each antibiotic within a diverse
drug panel, and identify potential future antibiotic resistant mutants. Our long-term goal is to develop an ENN-
based platform for automated transformation of data into axioms. Leveraging well-established collaborations
with colleagues of wide expertise, we will pursue these goals by combining our unique computational approaches
with experimental resources, including time-resolved protein aggregation assays, patient-derived tau fibrils co-
localized with sequence-specific RNA, high-throughput liquid culture antibiotic screens, multiplexed directed
evolution experiments of antibiotic resistance, and large in-house libraries of peptide and RNA mutant libraries.
This work lays the foundation for transforming large datasets into human-understandable rules
connecting sequence to function and relating these rules to physical mechanisms of structural dynamics. This in
turn could accelerate disease diagnosis and treatment.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金