ADMET evaluation in drug discovery: 15. Accurate prediction of rat oral acute toxicity using relevance vector machine and consensus modeling.

ADMET evaluation in drug discovery: 15. Accurate prediction of rat oral acute toxicity using relevance vector machine and consensus modeling.
复制标题

药物发现中的ADMET评估:15.利用相关向量机和共识模型准确预测大鼠口服急性毒性

DOI:
10.1186/s13321-016-0117-7
复制
发表时间:
2016
影响因子:
8.6
通讯作者:
Hou T
Hou T
中科院分区:
化学2区
文献类型:
--
作者:
Lei T;Li Y;Song Y;Li D;Sun H;Hou T

文献摘要

被引文献

相似文献

背景急性毒性的测定,以半数致死剂量(LD50)表示,是药物发现流程中最重要的步骤之一。由于哺乳动物口服急性毒性的体内检测耗时长、成本高,因此迫切需要开发口服急性毒性的计算机预测模型。结果基于包含7314种多种化学物质的大鼠口服LD50值的综合数据集,采用关联向量机(RVM)技术建立了口服急性毒性的回归模型,并与其他6种机器学习方法建立的回归模型进行了比较,这些回归模型包括K最近邻回归、随机森林(RF)、支持向量机、局部近似高斯过程、多层感知器集成和极端梯度增强。通过卡方统计选择原始分子描述符和结构指纹的子集(PubChem或SubFP)。由qext2对2 376个分子组成的测试集建立的各个定量构效关系模型的预测能力在0.572~0.659之间。结论综合考虑测试集的整体预测精度,建议采用拉普拉斯核和Rf的RVM建立对大鼠口服急性毒性具有较好预测能力的模型。通过结合各个模型的预测,开发了四个共识模型,为测试集产生了更好的预测能力(qext2=0.669-0.689)。最后,对一些与口服急性毒性相关的基本描述符和亚结构进行了识别和分析,它们可以作为避免毒性的属性或亚结构警报。我们认为,最优共识模型具有较高的预测精度,可以作为一种可靠的虚拟筛选工具来筛选出大鼠口服急性毒性较高的化合物。预测大鼠经口急性毒性的组合QSAR建模工作流程
BackgroundDetermination of acute toxicity, expressed as median lethal dose (LD50), is one of the most important steps in drug discovery pipeline. Because in vivo assays for oral acute toxicity in mammals are time-consuming and costly, there is thus an urgent need to develop in silico prediction models of oral acute toxicity.ResultsIn this study, based on a comprehensive data set containing 7314 diverse chemicals with rat oral LD50values, relevance vector machine (RVM) technique was employed to build the regression models for the prediction of oral acute toxicity in rate, which were compared with those built using other six machine learning approaches, includingk-nearest-neighbor regression, random forest (RF), support vector machine, local approximate Gaussian process, multilayer perceptron ensemble, and eXtreme gradient boosting. A subset of the original molecular descriptors and structural fingerprints (PubChem or SubFP) was chosen by the Chi squared statistics. The prediction capabilities of individual QSAR models, measured byqext2for the test set containing 2376 molecules, ranged from 0.572 to 0.659.ConclusionConsidering the overall prediction accuracy for the test set, RVM with Laplacian kernel and RF were recommended to build in silico models with better predictivity for rat oral acute toxicity. By combining the predictions from individual models, four consensus models were developed, yielding better prediction capabilities for the test set (qext2= 0.669–0.689). Finally, some essential descriptors and substructures relevant to oral acute toxicity were identified and analyzed, and they may be served as property or substructure alerts to avoid toxicity. We believe that the best consensus model with high prediction accuracy can be used as a reliable virtual screening tool to filter out compounds with high rat oral acute toxicity. Graphical abstractWorkflow of combinatorial QSAR modelling to predict rat oral acute toxicity