Graph kernels for chemical informatics

Graph kernels for chemical informatics
复制标题

DOI:
10.1016/j.neunet.2005.07.009
复制
发表时间:
2005-10-01
期刊:
影响因子:
7.8
通讯作者:
Baldi, P
Baldi, P
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ralaivola, L;Swamidass, SJ;Baldi, P

文献摘要

被引文献

相似文献

大型化合物存储库可用性的增加为机器学习方法应用于计算化学和化学信息学问题带来了新的挑战和机遇。由于化合物通常由其共价键的图形表示,因此该领域的机器学习方法必须能够处理大小可变的图形结构。在这里,我们首先简要回顾有关图内核的文献,然后介绍基于分子指纹思想的三个新内核(Tanimoto、MmMax、Hybrid),并使用从每个可能的顶点进行深度优先搜索来计算深度高达 d 的标记路径。这些内核应用于三个分类问题,以预测三个公开数据集的致突变性、毒性和抗癌活性。这些内核的性能至少与之前文献报道的性能相当,而且通常优于之前报道的性能,在 Mutag 数据集上达到 91.5% 的准确率,在 PTC(预测毒理学挑战)数据集上达到 65-67%,在 NCI(国家癌症研究所)数据集上达到 72%。简要讨论了这些内核以及其他利用分子 1D 或 3D 表示的内核的属性和权衡。 (c) 2005 Elsevier Ltd. 保留所有权利。
Increased availability of large repositories of chemical compounds is creating new challenges and opportunities for the application of machine learning methods to problems in computational chemistry and chemical informatics. Because chemical compounds are often represented by the graph of their covalent bonds, machine learning methods in this domain Must be capable of processing graphical structures with variable size. Here, we first briefly review the literature on graph kernels and then introduce three new kernels (Tanimoto, MmMax, Hybrid) based on the idea of molecular fingerprints and counting labeled paths of depth up to d using depth-first search from each possible vertex. The kernels are applied to three classification problems to predict mutagenicity, toxicity, and anti-cancer activity on three publicly available data sets. The kernels achieve performances at least comparable, and most often superior, to those previously reported in the literature reaching accuracies of 91.5% on the Mutag dataset, 65-67% on the PTC (Predictive Toxicology Challenge) dataset, and 72% on the NCI (National Cancer Institute) dataset. Properties and tradeoffs of these kernels, as well as other proposed kernels that leverage 1D or 3D representations of molecules, are briefly discussed. (c) 2005 Elsevier Ltd. All rights reserved.