融合分子向量图卷积等方法的高通量药物筛选及其在新冠病毒RdRp上的运用
批准号:
62106253
项目类别:
青年科学基金项目(C类)
资助金额:
30.0 万元
负责人:
张海平
依托单位:
学科分类:
交叉学科驱动的人工智能
结题年份:
2024
批准年份:
2021
项目状态:
已结题
项目参与者:
张海平
中文摘要
高效准确的计算机辅助药物高通量逐级筛选管道是加速全新药物开发的关键。由于类药可合成小分子的数量庞大,蛋白配体复合物难以被精确获取,导致速度过低,精度不足。目前已有的蛋白-小分子预测模型往往缺乏训练负极,难以有效地表示关键信息,且数据量不足,过于依赖复合物结构,缺乏不同方法有效联合。为解决高速筛选瓶颈,本项目拟将蛋白质的口袋信息和小分子转化成分子向量信息,并引入图卷积神经网络保存关键空间信息,运用密集深度全连接网络进行学习,产生的模型具有精度高,速度快等优点,有望快速准确地将大数据库中的活性小分子富集到小数据集。此外,拟开发多个基于结构的深度学习模型,与分子动力学模拟等技术一起用于精细筛选。最终,拟联合多种方法,用于针对新冠病毒RdRp靶点的千万级别超大规模逐级药物虚拟筛选和设计。本研究将为药物筛选和设计提供关键理论和方法支持,有望应用于新冠肺炎等疾病的药物开发。
英文摘要
Developing efficient and accurate computational-aided extremely large-scale virtual screening platforms would help facilitate novel drug development. Because the diversity of feasible synthesizable drug-like compounds is extremely large, traditional computational methods are limited in fulfilling current requirements. How to fast and accurately enrich the activity compounds into the smaller dataset of top score prediction is the major challenge. We convert the protein active cavity and ligand into a molecular vector representation, using the densely fully-connected neural network and Graph Convolutional Network(GCN) as model architectures. The model allows us to avoid the difficult challenge of obtaining accurate protein-ligand structure complexes. This model is fast and easy to use. It is suitable for large-scale target searching for certain ligands, and is suitable for large-scale virtual screening for lead drugs over drug-like molecule databases with millions of small molecules. On the other hand, we develop several deep learning based protein-ligand prediction models for later stage middle and small-scale virtual screening. The pocket molecular dynamics novel method was also developed and incorporated into the platform for later stage elaborate selection of binding compounds. In this research, we will systematically evaluate the model accuracy on many available experiment data. We also will check the enrichment ability for active compounds in a large compound database, hence find out the suitable type of proteins for its application. We will use this pipeline in step by step virtual screening of compounds for RdRp,which is an important therapeutic target for most RNA virus. Our research will provide key methods and new tools for drug virtual screening and design, and are promising in promoting drug development against SARS-CoV-2.
高效准确的评估蛋白-小分子是否作用、作用强度,对大规模筛选针对靶点药物有重要意义。配合全新化合物生成模型,有望加速全新药物开发。而以注意力机制,图卷积等为代表的深度学习由于自动高效提取数据特征,能保留物理化学性质及空间特征,已经在精度和效率上表现出了明显优势。目前已有的蛋白-小分子预测模型往往依赖复合物信息,速度和准确度不足,为解决药物高速筛选瓶颈,本项目收集已有蛋白-小分子 相互作用实验数据,同时将蛋白-小分子的相互作用关键界面信息作为输入,并引入统计势能,预训练向量,图卷积神经网络,注意力机制等技术用以保存关键空间及物理化学信息,最终开发多套互补模型用于预测蛋白-小分子是否结合,亲和力预测,并与计算生物相关技术一起组成筛选管道。最终,我们将开发方法和管道成功用于针对RdRp,ns4b, TIPE2/3,GPR35,CHD1L等抗病毒或抗癌靶点的活性分子筛选设计及实验验证,均成功发现活性化合物,大部分成功率超过20%。说明方法的有效性。类似策略我们也成功迁移到了抗体设计和疫苗设计,显示出优良表现。此外开发亲和力测试模型在CASP16比赛中获得了小分子组亲和力测试赛道第一名,说明该方法已经达到国际一流水平。课题支持期间发表相关第一作或通讯论文十多篇,且绝大部分是JCR一区。
国内基金
海外基金