Predicting human liver microsomal stability with machine learning techniques

Predicting human liver microsomal stability with machine learning techniques
复制标题

DOI:
10.1016/j.jmgm.2007.06.005
复制
发表时间:
2008-02-01
影响因子:
2.9
通讯作者:
Honma, Teruki
Honma, Teruki
中科院分区:
生物学4区
文献类型:
--
作者:
Sakiyama, Yojiro;Yuki, Hitomi;Honma, Teruki

文献摘要

被引文献

相似文献

为了确保药物研究的持续管道,主要候选人必须在药物发现过程中具有适当的代谢稳定性。体外ADMET(吸收、分布、代谢、消除和毒性)筛选为我们提供了有关化合物代谢稳定性的有用信息。然而,在合成阶段之前,需要一个有效的过程,以处理来自大型化合物库和高通量筛选的大量数据。在这里,我们通过各种计算机机器学习,如随机森林,支持向量机(SVM),逻辑回归和递归分区,推导出内部化合物数据集的化学结构及其代谢稳定性之间的关系。为了建立模型,使用了1952种专利化合物,包括两个类别(稳定/不稳定)和193个由分子操作环境计算的描述符。使用测试化合物的结果已经证明,所有分类器都产生了令人满意的结果(准确度> 0.8,灵敏度> 0.9,特异性> 0.6,并且精确度> 0.8)。最重要的是,通过随机森林以及SVM的分类在独立验证集中产生了约0.7的Kappa值,略高于其他分类工具。这些结果表明,非线性/集成为基础的分类方法可能证明是有用的,在硅片ADME建模领域。(C)2007年爱思唯尔公司All rights reserved.
To ensure a continuing pipeline in pharmaceutical research, lead candidates must possess appropriate metabolic stability in the drug discovery process. In vitro ADMET (absorption, distribution, metabolism, elimination, and toxicity) screening provides us with useful information regarding the metabolic stability of compounds. However, before the synthesis stage, an efficient process is required in order to deal with the vast quantity of data from large compound libraries and high-throughput screening. Here we have derived a relationship between the chemical structure and its metabolic stability for a data set of in-house compounds by means of various in silico machine learning such as random forest, support vector machine (SVM), logistic regression, and recursive partitioning. For model building, 1952 proprietary compounds comprising two classes (stable/unstable) were used with 193 descriptors calculated by Molecular Operating Environment. The results using test compounds have demonstrated that all classifiers yielded satisfactory results (accuracy > 0.8, sensitivity > 0.9, specificity > 0.6, and precision > 0.8). Above all, classification by random forest as well as SVM yielded kappa values of approximately 0.7 in an independent validation set, slightly higher than other classification tools. These results suggest that nonlinear/ensemble-based classification methods might prove useful in the area of in silico ADME modeling. (C) 2007 Elsevier Inc. All rights reserved.