Accelerating the Screening of Small Peptide Ligands by Combining Peptide-Protein Docking and Machine Learning.

Accelerating the Screening of Small Peptide Ligands by Combining Peptide-Protein Docking and Machine Learning.
复制标题

DOI:
10.3390/ijms241512144
复制
发表时间:
2023-07-29
影响因子:
5.6
通讯作者:
Daunert, Sylvia
Daunert, Sylvia
中科院分区:
生物学2区
文献类型:
--
作者:
Codina, Josep-Ramon;Mascini, Marcello;Dikici, Emre;Deo, Sapna K.;Daunert, Sylvia

文献摘要

参考文献

相似文献

本研究提出了一种结合机器学习(ML)和分子对接的新型管道,通过预测肽-蛋白对接来加速小肽配体筛选过程。分析了8种ML算法的潜力。值得注意的是,光梯度增强机(LightGBM)尽管具有与同类产品相当的f1分数和精度,但显示出卓越的计算效率。使用LightGBM对16万个肽配体的整个四肽库与四种病毒包膜蛋白的肽-蛋白对接性能进行分类。图书馆被分为两组,“表现较好”和“表现较差”。通过仅在1%的四肽库上训练LightGBM算法,我们成功地对剩余的99%进行了分类,准确率范围为0.81-0.85,f1分数在0.58-0.67之间。使用了三种不同的分子对接软件来证明该过程不依赖于软件。通过可调的概率阈值(从0.5到0.95),该过程可以加速至少10倍,并且与不使用ML的方法相比仍可获得90-95%的并发性。本研究验证了机器学习与分子对接在不依赖高性能计算能力的情况下快速识别顶级肽的效率,使其成为筛选潜在生物活性化合物的有效工具。
This research introduces a novel pipeline that couples machine learning (ML), and molecular docking for accelerating the process of small peptide ligand screening through the prediction of peptide-protein docking. Eight ML algorithms were analyzed for their potential. Notably, Light Gradient Boosting Machine (LightGBM), despite having comparable F1-score and accuracy to its counterparts, showcased superior computational efficiency. LightGBM was used to classify peptide-protein docking performance of the entire tetrapeptide library of 160,000 peptide ligands against four viral envelope proteins. The library was classified into two groups, ‘better performers’ and ‘worse performers’. By training the LightGBM algorithm on just 1% of the tetrapeptide library, we successfully classified the remaining 99%with an accuracy range of 0.81–0.85 and an F1-score between 0.58–0.67. Three different molecular docking software were used to prove that the process is not software dependent. With an adjustable probability threshold (from 0.5 to 0.95), the process could be accelerated by a factor of at least 10-fold and still get 90–95% concurrence with the method without ML. This study validates the efficiency of machine learning coupled to molecular docking in rapidly identifying top peptides without relying on high-performance computing power, making it an effective tool for screening potential bioactive compounds.
DOI: 10.3390/biom9090498
发表时间: 2019-09-01
期刊: BIOMOLECULES
影响因子: 5.5
作者:
Mascini, Marcello;Dikici, Emre;Daunert, Sylvia
通讯作者: Daunert, Sylvia
DOI: 10.1002/bip.20296
发表时间: 2005-01-01
期刊: BIOPOLYMERS
影响因子: 2.9
作者:
Mei, H;Liao, ZH;Li, SZ
通讯作者: Li, SZ
DOI: 10.1093/bioinformatics/btx005
发表时间: 2017-05-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Hou, Qingzhen;De Geest, Paul F. G.;Feenstra, K. Anton
通讯作者: Feenstra, K. Anton
DOI: 10.1214/aos/1013203451
发表时间: 2001-10-01
影响因子: 4.5
作者:
Friedman, JH
通讯作者: Friedman, JH
DOI: 10.1111/j.1747-0285.2008.00641.x
发表时间: 2008-04-01
影响因子: 3
作者:
Liang, Guizhao;Chen, Guohua;Li, Zhiliang
通讯作者: Li, Zhiliang