A novel logic-based approach for quantitative toxicology prediction

A novel logic-based approach for quantitative toxicology prediction
复制标题

DOI:
10.1021/ci600223d
复制
发表时间:
2007-05-01
影响因子:
5.6
通讯作者:
Sternberg, Michael J. E.
Sternberg, Michael J. E.
中科院分区:
化学2区
文献类型:
--
作者:
Amini, Ata;Muggleton, Stephen H.;Sternberg, Michael J. E.

文献摘要

被引文献

相似文献

迫切需要准确的计算机方法来预测被引入环境或正在开发成新药物的分子的毒性。预测毒理学属于结构活性关系 (SAR) 领域,并且已使用许多方法来推导此类 SAR。先前的工作表明,归纳逻辑编程(ILP)是一种强大的方法,可以克服其他一些 SAR 方法面临的几个主要困难,例如分子叠加。 ILP 方法在关系框架内利用化学子结构进行推理,并产生化学上可理解的规则。在这里,我们报告了一种通用的新方法,即支持向量归纳逻辑编程(SVILP),它将本质上基于 ILP 的定性 SAR 扩展到定量建模。首先,ILP 用于学习规则,然后在新颖的内核中使用规则的预测来导出支持向量泛化模型。对于具有已知黑头鲦鱼毒性的 576 个分子的高度异质数据集,化学描述符方法 (CHEM) 和 SVILP 的交叉验证相关系数 (R-CV(2)) 分别为 0.52 和 0.66。 ILP、CHEM 和 SVILP 方法分别正确预测了 55%、58% 和 73% 的有毒分子。在一组 165 个看不见的分子中,商业软件 TOPKAT 和 SVILP 的 R-2 值分别为 0.26 和 0.57。在所有计算中,与其他方法相比,SVILP 都显示出显着的改进。 SVILP 方法的一个主要优点是,它自动且一致地使用 ILP 来导出规则(大部分是新颖的),描述毒性警报的片段。 SVILP 是一种通用的机器学习方法,有潜力解决与化学信息学相关的许多问题,包括计算机药物设计中的问题。
There is a pressing need for accurate in silico methods to predict the toxicity of molecules that are being introduced into the environment or are being developed into new pharmaceuticals. Predictive toxicology is in the realm of structure activity relationships ( SAR), and many approaches have been used to derive such SAR. Previous work has shown that inductive logic programming (ILP) is a powerful approach that circumvents several major difficulties, such as molecular superposition, faced by some other SAR methods. The ILP approach reasons with chemical substructures within a relational framework and yields chemically understandable rules. Here, we report a general new approach, support vector inductive logic programming (SVILP), which extends the essentially qualitative ILP-based SAR to quantitative modeling. First, ILP is used to learn rules, the predictions of which are then used within a novel kernel to derive a support-vector generalization model. For a highly heterogeneous dataset of 576 molecules with known fathead minnow fish toxicity, the cross-validated correlation coefficients (R-CV(2)) from a chemical descriptor method ( CHEM) and SVILP are 0.52 and 0.66, respectively. The ILP, CHEM, and SVILP approaches correctly predict 55, 58, and 73%, respectively, of toxic molecules. In a set of 165 unseen molecules, the R-2 values from the commercial software TOPKAT and SVILP are 0.26 and 0.57, respectively. In all calculations, SVILP showed significant improvements in comparison with the other methods. The SVILP approach has a major advantage in that it uses ILP automatically and consistently to derive rules, mostly novel, describing fragments that are toxicity alerts. The SVILP is a general machine-learning approach and has the potential of tackling many problems relevant to chemoinformatics including in silico drug design.