课题基金 / 基金详情

Improving the Accuracy of Implicit Solvents with a Physics-Guided Neural Network

Improving the Accuracy of Implicit Solvents with a Physics-Guided Neural Network
利用物理引导神经网络提高隐式溶剂的准确性
批准号:
10669809
负责人:
Negin Forouzesh
金额:
$18.25万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-08-01 至 2026-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
项目摘要/摘要 药物发现是生物科学中最具挑战性的任务之一;它需要大约10-15年的时间和 发现一种新药平均需要20亿美元。药物发现的主要目标是识别类似药物的COM- 能够调节特定fic生物靶标(蛋白质)的磅(配体)。蛋白质的一个重要特征--配体 相互作用是发生在蛋白质和配体之间的结合自由能变化G。 配基的依恋。这一物理化学特征在很大程度上决定了蛋白质和配体相互作用的强度 对药物设计的理解尤其有用。虽然湿实验室实验准确地估计了G,但他们 是明显的缓慢,昂贵,和费力的fi。另一方面,计算模拟可以使显著的fi 更快地估计G,并阐明各种结构的结合机制,这些结构本来可以 以其他方式进行检查是很复杂的。隐含的溶剂框架,将溶剂视为一个连续体 与水的介电和非极性性质相比,提供了更有效的G的fi估计 到其他计算方法,如炼金术自由能方法。尽管取得了显著的进展 在隐式溶剂建模中,对其准确性的严重担忧源于潜在的物理- Cal近似。这项研究将使用现代机器学习技术来弥合精度差距 在G计算方面,基于物理的隐式溶剂模型与实验参考之间的关系。在……里面 特别是,实验数据将被集成到广义Born(GB)隐式溶剂模型中,以便与 坚持物理模型,新的结构特征可以提高精度。除了模型之外, 准确性,必须保持可解释性(这是模型简单性的原因)和可转移性(即 确保在不同数据集上的一致性能)。为此,一种新的多目标损失函数将是 引入了考虑“准确性”、“可解释性”和“可转移性”的概念。标准蛋白质-配基 数据库、基准测试和数据集将用于设计提议的混合模型,包括主机-来宾 系统、Sampl Challenges基准、PDBind和BindingDB。虽然其中一些来源包含干净的 数据,许多需要进一步的后处理,为运行国标模式做准备。仔细的数据准备将 通过遵循标准协议并通过流行的Web服务来执行。它的模块化特征 提议的物理数据模型将允许测试隐含溶剂的各种fl偏好(基于物理的模型)和 Modifi阳离子到建议的图卷积网络(数据驱动模型)。混合动力车的这一fl灵活性 模型促进了基于经典物理和现代数据驱动的新的跨学科研究 结束了。fiNAL源代码和参数化数据集将免费向公众提供。他们可能是 在药物发现的早期阶段结合到候选药物的高通量虚拟筛选中。 这项研究的结果将为fi生物分子模型界提供一种构建 用于研究蛋白质-配体相互作用的新颖、准确和有效的fi计算模型。
英文摘要
Project Summary/Abstract Drug discovery is one of the most challenging tasks in biological sciences; it takes about 10-15 years and $2 billion on average to discover a new drug. The main goal in drug discovery is identifying drug-like com- pounds (ligands) capable of modulating specific biological targets (proteins). One key feature of protein-ligand interactions is the binding free energy change, G, that occurs between the protein and the ligand upon the ligand's attachment. This physiochemical feature heavily dictates how strongly a protein and ligand interact and is particularly useful to understand for drug design. While wet-lab experiments accurately estimate G, they are significantly slow, costly, and laborious. On the other hand, computational simulations enable significantly faster estimation of G and shed light on the binding mechanism of various structures that could have been complicated to be examined otherwise. The implicit solvent framework, which treats solvent as a continuum with the dielectric and non-polar properties of water, offer much more efficient estimation of G compared to other computational methodologies, such as alchemical free energy methods. Despite noticeable progress in implicit solvent modeling, serious concerns about its accuracy remain that stem from the underlying physi- cal approximations. This research will employ modern machine learning techniques to bridge the accuracy gap between a physics-based implicit solvent model and experimental references in terms of G calculations. In particular, experimental data will be integrated into a generalized Born (GB) implicit solvent model so that with adherence to the physical model, new structural features could improve the accuracy. In addition to the model accuracy, it is essential to retain interpretability (that accounts for the model simplicity) and transferability (that assures consistent performance on different datasets). To this end, a novel multi-objective loss function will be introduced that takes “accuracy”, “interpretability”, and “transferability” into consideration. Standard protein-ligand databases, benchmarks, and datasets will be used for designing the proposed hybrid model, including host-guest systems, SAMPL challenge benchmarks, PDBbind, and BindingDB. While some of these sources contain clean data, many require further post-processing to prepare for running the GB model. Careful data preparation will be performed by following standard protocols and via popular web services. The modular characteristics of the proposed physics-data model will allow for testing various flavors of implicit solvent (physics-based model) and modifications to the proposed Graph Convolutional Network (data-driven model). This flexibility of the hybrid model facilitates new interdisciplinary research between the classical physics-based and the modern data-driven ends. The final source code and parameterized datasets will be available freely to the public. They could be incorporated into the high-throughput virtual screening of candidate drugs in the early stages of drug discovery. The outcome of this research will benefit the biomolecular modeling community by providing an approach to build novel, accurate, and efficient computational models for studying protein-ligand interactions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金