QSAR-derived affinity fingerprints (part 2): modeling performance for potency prediction

QSAR-derived affinity fingerprints (part 2): modeling performance for potency prediction
复制标题

DOI:
10.1186/s13321-020-00444-5
复制
发表时间:
2020-06-05
影响因子:
8.6
通讯作者:
Svozil, Daniel
Svozil, Daniel
中科院分区:
化学2区
文献类型:
--
作者:
Cortes-Ciriano, Isidro;Skuta, Ctibor;Svozil, Daniel

文献摘要

被引文献

相似文献

亲和指纹通过一系列分析报告小分子的活性,从而允许收集关于结构不同化合物的生物活性的信息,其中仅基于化学结构的模型往往是有限的,并对复杂的生物终点进行建模,如人体毒性和体外癌细胞系敏感性。在这里,我们建议使用计算预测的生物活性曲线作为化合物描述符来模拟化合物的体外活性。为此,我们应用并验证了一个用于计算QSAR衍生亲和指纹(QAFFP)的框架,该框架使用了一组1360个QSAR模型,这些模型使用来自CHEMBL数据库的K-I、K-d、IC50和EC50数据。因此,QAFFP代表了一种基于化合物在生物活性空间中的相似性来编码和关联化合物的方法。为了对QAFFP的预测能力进行基准测试,我们从ChEMBL数据库中收集了18个广泛用于临床前药物发现的不同癌细胞株的IC50数据,以及25个不同的蛋白质靶标数据集。这项研究是对第一部分的补充,在第一部分中,评估了QAFFP在相似性搜索、支架跳跃和生物活性分类方面的性能。尽管存在固有的噪声,但我们表明,使用QAFFP作为描述符会导致在类似于0.65-0.95PIC(50)单位范围内对测试集的预测误差,这与CHEMBL(0.76-1.00PIC(50)单位)中生物活性数据的估计不确定性相当。我们发现,QAFFP的预测能力略逊于摩根2指纹以及一维和二维物理化学描述符,其影响大小在0.02-0.08PIC(50)单位范围内。在QAFFP的生成中加入预测能力较低的QSAR模型并不会导致预测能力的提高。鉴于我们用来计算QAFFP的QSAR模型是仅根据数据可用性来选择的,我们预计使用更多不同和具有生物意义的目标生成的QAFFP将获得更好的建模结果。数据集和Python代码可在上公开获得。
Affinity fingerprints report the activity of small molecules across a set of assays, and thus permit to gather information about the bioactivities of structurally dissimilar compounds, where models based on chemical structure alone are often limited, and model complex biological endpoints, such as human toxicity and in vitro cancer cell line sensitivity. Here, we propose to model in vitro compound activity using computationally predicted bioactivity profiles as compound descriptors. To this aim, we apply and validate a framework for the calculation of QSAR-derived affinity fingerprints (QAFFP) using a set of 1360 QSAR models generated using K-i, K-d, IC50 and EC50 data from ChEMBL database. QAFFP thus represent a method to encode and relate compounds on the basis of their similarity in bioactivity space. To benchmark the predictive power of QAFFP we assembled IC50 data from ChEMBL database for 18 diverse cancer cell lines widely used in preclinical drug discovery, and 25 diverse protein target data sets. This study complements part 1 where the performance of QAFFP in similarity searching, scaffold hopping, and bioactivity classification is evaluated. Despite being inherently noisy, we show that using QAFFP as descriptors leads to errors in prediction on the test set in the similar to 0.65-0.95 pIC(50) units range, which are comparable to the estimated uncertainty of bioactivity data in ChEMBL (0.76-1.00 pIC(50) units). We find that the predictive power of QAFFP is slightly worse than that of Morgan2 fingerprints and 1D and 2D physicochemical descriptors, with an effect size in the 0.02-0.08 pIC(50) units range. Including QSAR models with low predictive power in the generation of QAFFP does not lead to improved predictive power. Given that the QSAR models we used to compute the QAFFP were selected on the basis of data availability alone, we anticipate better modeling results for QAFFP generated using more diverse and biologically meaningful targets. Data sets and Python code are publicly available at.