Predicting protein-ligand interactions based on bow-pharmacological space and Bayesian additive regression trees

Predicting protein-ligand interactions based on bow-pharmacological space and Bayesian additive regression trees
复制标题

DOI:
10.1038/s41598-019-43125-6
复制
发表时间:
2019-05-22
期刊:
影响因子:
4.6
通讯作者:
Wei, Dong-Qing
Wei, Dong-Qing
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Li, Li;Koh, Ching Chiek;Wei, Dong-Qing

文献摘要

被引文献

相似文献

识别潜在的蛋白质-配体相互作用是药物发现领域的核心,因为它有助于识别潜在的新药先导,有助于从命中到先导的进展,预测批准的药物或候选物的副作用的潜在脱靶解释,以及去孤儿表型命中。为了快速识别蛋白质-配体相互作用,我们在这里提出了一种新的化学基因组学算法,用于预测蛋白质-配体相互作用,使用一种新的机器学习方法和新的描述符类。该算法适用于贝叶斯加性回归树(BART)上新提出的蛋白质化学空间,称为弓药理空间。该空间跨越三个不同的子空间,覆盖蛋白质空间,配体空间和相互作用空间。因此,该模型扩展了依赖于这些子空间中的一个或两个的经典靶标预测或化学基因组学建模的范围。我们的模型表现出出色的预测能力,在对四个人类靶数据集(包括酶、核受体、离子通道和G蛋白偶联受体)进行评估时,准确率高达94.5-98.4%。BART提供了蛋白质和配体之间相互作用的可能性的可靠概率描述,其可用于在小分子开发的发现和警戒阶段进行的测定的优先级排序。
Identifying potential protein-ligand interactions is central to the field of drug discovery as it facilitates the identification of potential novel drug leads, contributes to advancement from hits to leads, predicts potential off-target explanations for side effects of approved drugs or candidates, as well as de-orphans phenotypic hits. For the rapid identification of protein-ligand interactions, we here present a novel chemogenomics algorithm for the prediction of protein-ligand interactions using a new machine learning approach and novel class of descriptor. The algorithm applies Bayesian Additive Regression Trees (BART) on a newly proposed proteochemical space, termed the bow-pharmacological space. The space spans three distinctive sub-spaces that cover the protein space, the ligand space, and the interaction space. Thereby, the model extends the scope of classical target prediction or chemogenomic modelling that relies on one or two of these subspaces. Our model demonstrated excellent prediction power, reaching accuracies of up to 94.5-98.4% when evaluated on four human target datasets constituting enzymes, nuclear receptors, ion channels, and G-protein-coupled receptors . BART provided a reliable probabilistic description of the likelihood of interaction between proteins and ligands, which can be used in the prioritization of assays to be performed in both discovery and vigilance phases of small molecule development.