Network-based piecewise linear regression for QSAR modelling

Network-based piecewise linear regression for QSAR modelling
复制标题

DOI:
10.1007/s10822-019-00228-6
复制
发表时间:
2019-10-18
影响因子:
3.5
通讯作者:
Tsoka, Sophia
Tsoka, Sophia
中科院分区:
生物学3区
文献类型:
--
作者:
Cardoso-Silva, Jonathan;Papageorgiou, Lazaros G.;Tsoka, Sophia

文献摘要

被引文献

相似文献

定量构效关系(QSAR)模型在药物发现的各个领域至关重要,例如在先导优化和虚拟筛选中。最近,对既能预测又能解释的模型的需求得到了强调。本文提出了一种将网络分析与分段线性回归相结合构建可解释QSAR模型的新方法。提出的算法modSAR采用两步分割数据。首先,与共同目标相关的化合物根据其结构相似性表示为网络,揭示具有相似化学性质的模块。其次,将每个模块细分为子集(区域),每个子集由一个独立的线性方程建模。研究人员对从ChEMBL获得的5个蛋白质抑制剂数据集的QSAR模型进行了比较分析,结果表明,modSAR模型的预测精度与随机森林和支持向量机等流行算法相似。此外,我们证明了modSAR建立的模型是可解释的,能够评估化合物的适用范围,并很好地服务于虚拟筛选和新药先导开发等任务。
Quantitative Structure-Activity Relationship (QSAR) models are critical in various areas of drug discovery, for example in lead optimisation and virtual screening. Recently, the need for models that are not only predictive but also interpretable has been highlighted. In this paper, a new methodology is proposed to build interpretable QSAR models by combining elements of network analysis and piecewise linear regression. The algorithm presented, modSAR, splits data using a two-step procedure. First, compounds associated with a common target are represented as a network in terms of their structural similarity, revealing modules of similar chemical properties. Second, each module is subdivided into subsets (regions), each of which is modelled by an independent linear equation. Comparative analysis of QSAR models across five data sets of protein inhibitors obtained from ChEMBL is reported and it is shown that modSAR offers similar predictive accuracy to popular algorithms, such as Random Forest and Support Vector Machine. Moreover, we show that models built by modSAR are interpretatable, capable of evaluating the applicability domain of the compounds and serve well tasks such as virtual screening and the development of new drug leads.