Geometric Deep Learning for Binding Affinity Prediction
Geometric Deep Learning for Binding Affinity Prediction
批准号:
2597682
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
药物发现是一个非常昂贵和耗时的过程。将一种药物推向市场的平均成本是13亿美元,每九年翻一番,需要十到十五年。机器和深度学习技术的发展有望提高药物发现过程的效率,并阻止生产力的下降。基于结构的药物发现是发现方法之一,它使用计算方法和蛋白质的3D结构来识别与靶点结合的新型药物。在过去的十年中,已经使用机器学习开发了可以预测蛋白质和小分子药物之间结合亲和力的准确评分函数。这些基于机器学习(ML)的评分函数比现有方法提高了准确性。然而,它们主要是使用结合蛋白质-配体复合物的解析晶体结构进行训练和验证的。这并不能准确地代表现实世界的药物发现场景,其中结合的蛋白质-配体复合物的晶体结构不可用。有时可能根本没有蛋白质的晶体结构,必须使用模型结构。此外,基于ML的评分函数通常使用来自单一来源的数据进行训练和验证,并且通常无法推广到新的数据集。为了区分由训练数据中的偏差导致的评分函数的预测和那些没有偏差的预测,必须完全理解模型的推理。当前的评分功能不提供该信息,因此在药物发现中使用的领域中具有降低的信任。归因等方法在阐明为什么评分函数做出了某些预测方面显示出了希望。归因可以用来检查预测是否基于蛋白质和药物之间的相互作用,而不仅仅是药物的组成,这是评分功能的常见陷阱。本DPhil旨在使用深度学习开发一个绑定亲和力模型,该模型已明确设计用于解决该领域现有模型中的缺陷。几何深度学习架构(如等变图神经网络)已成功用于此问题,并与归因相结合,表明评分函数确实学习生物分子相互作用,而不是偏见来进行预测。这些架构将建立在扩展建模结构的准确性,而不仅仅是实验确定的晶体结构,以允许增加评分功能的实际应用的可靠性。一个准确、可靠、易于解释的评分函数在任何药物发现项目中都是有价值的,并将作为开源软件提供给所有人。该研究将与医学研究慈善机构LifeArc及其员工Andy Merritt博士和Kristian Birchall博士合作完成,他们在加速医疗创新方面拥有宝贵的经验,为患者取得突破。该项目福尔斯EPSRC AI和数据科学工程,健康和政府(ASG)研究领域。
英文摘要
Drug discovery is an incredibly expensive and time-consuming process. The average cost of bringing a drug to market is $1.3 billion, which is doubling every nine years, and takes ten to fifteen years. The development of machine and deep learning techniques have promised to improve the efficiency of the drug discovery process and arrest this decline in productivity. Structure-based drug discovery, one method of discovery, uses computational methods and the 3D structure of a protein to identify novel drugs that bind to the target. Accurate scoring functions that can predict the binding affinity between a protein and a small molecule drug have been developed using machine learning over the last decade. These machine learning (ML)-based scoring functions have improved accuracy over pre-existing methods. However, they have been primarily trained and validated using solved crystal structures of bound protein-ligand complexes. This does not accurately represent a real-world drug discovery scenario where crystal structures for the bound protein-ligand complexes are not available. Sometimes there may be no crystal structure of the protein at all, and modelled structures must be used. In addition, ML-based scoring functions are typically trained and validated using data from a single source and often fail to generalise to novel data sets. To discern between predictions from the scoring functions that result from a bias in the training data and those that do not, the model's reasoning must be fully understood. Current scoring functions do not provide this information and so have a reduced trust in the field to be used in drug discovery. Approaches such as attribution have shown promise in elucidating why the scoring function has made certain predictions. Attribution can be used to check whether predictions have been made based on the interactions between the protein and drug instead of just on the composition of the drug alone, a common pitfall for scoring functions. This DPhil aims to develop a binding affinity model using deep learning that has been explicitly designed to address the flaws currently in existing models in the field. Geometric deep learning architectures, such as equivariant graph neural networks, have been successfully utilised for this problem and combined with attribution to show that the scoring function does learn biomolecular interactions instead of bias to make predictions. These architectures will be built upon to extend the accuracy for modelled structures, not just experimentally determined crystal structures, to allow increased reliability for realistic applications of scoring functions. A scoring function that is accurate, reliable, and easily interpretable will be valuable in any drug discovery project and will be available for all as open-source software. The research will be completed in collaboration with the medical research charity LifeArc and their employees Dr Andy Merritt and Dr Kristian Birchall, who have valuable experience in accelerating healthcare innovation to make breakthroughs for patients. The project falls within the EPSRC AI and Data Science for Engineering, Health, and Government (ASG) research area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
-
批准号:2026JJ81909
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:胡曦
-
依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
-
批准号:12271434
-
项目类别:面上项目
-
资助金额:46万元
-
批准年份:2022
-
负责人:贺小伟
-
依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
-
批准号:2020A151501709
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2020
-
负责人:谢怡
-
依托单位:
面向Deep Web的数据整合关键技术研究
-
批准号:61872168
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2018
-
负责人:董永权
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于语义计算的海量Deep Web知识探索机制研究
-
批准号:61272411
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:赵峰
-
依托单位:
Deep Web数据集成查询结果抽取与整合关键技术研究
-
批准号:61100167
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:董永权
-
依托单位:
面向Deep Web的大规模知识库自动构建方法研究
-
批准号:61170020
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2011
-
负责人:崔志明
-
依托单位:
Deep Web敏感聚合信息保护方法研究
-
批准号:61003054
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2010
-
负责人:赵朋朋
-
依托单位: