Similarity searching of chemical databases using atom environment descriptors (MOLPRINT 2D): Evaluation of performance

Similarity searching of chemical databases using atom environment descriptors (MOLPRINT 2D): Evaluation of performance
复制标题

DOI:
10.1021/ci0498719
复制
发表时间:
2004-09-01
期刊:
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子:
--
通讯作者:
Reiling, S
Reiling, S
中科院分区:
其他
文献类型:
--
作者:
Bender, A;Mussa, HY;Reiling, S

文献摘要

被引文献

相似文献

将一种基于原子环境、基于信息增益的特征选择和朴素贝叶斯分类器的分子相似性搜索技术应用于一系列不同的数据集,并与其他搜索方法进行了性能比较。原子环境是存在于分子结构中每个重原子的拓扑距离上的重原子的计数向量。在这个应用程序中,使用MDL药物数据报告数据库中最近发布的超过100000个分子的数据集,原子环境方法似乎优于排名分数融合和二进制核识别,这两者都与Unity指纹结合使用。排名前5%的文库的总体检索率比排名第二的统一指纹和二值核鉴别方法的检索率高出近10%(相对数字高出14%以上)。在11组活性化合物中,有10组的原子环境与朴素贝叶斯分类器相结合是较好的方法,而在其余数据集中,数据融合和二值核识别结合Unity指纹是首选方法。二进制内核识别结合Unity指纹通常在整体性能上排名第二。性能上的差异很大程度上归因于所使用的不同分子描述符。如果将这些描述符与谷本系数的组合进行比较,Atom环境的性能将大大优于Unity指纹。朴素贝叶斯分类器结合基于信息增益的特征选择和合理数量特征的选择,在实验中表现出与二值核鉴别相同的性能,并对这些分类方法进行了比较。当在单氨基氧化酶数据集上使用时,原子环境和朴素贝叶斯分类器在训练和测试化合物50/50分割的情况下表现得和二元核区分一样好。在稀疏训练数据的情况下,发现二值核判别在这个特定的数据集上是优越的。在第三个数据集上,原子环境描述符与谷本相似系数结合使用时,显示出比这里测试的其他2D指纹更高的检索率。特征选择是决定算法性能的关键步骤。分子的原子环境表示被发现比Unity指纹更有效的生物受体相似性计算类型在这里检查。结合评分前的信息和包括非活性化合物的信息,如贝叶斯分类器和二值核判别,被发现优于后验数据融合(在这里测试的数据集中)。
A molecular similarity searching technique based on atom environments, information-gain-based feature selection, and the naive Bayesian classifier has been applied to a series of diverse datasets and its performance compared to those of alternative searching methods. Atom environments are count vectors of heavy atoms present at a topological distance from each heavy atom of a molecular structure. In this application, using a recently published dataset of more than 100000 molecules from the MDL Drug Data Report database, the atom environment approach appears to outperform fusion of ranking scores as well as binary kernel discrimination, which are both used in combination with Unity fingerprints. Overall retrieval rates among the top 5% of the sorted library are nearly 10% better (more than 14% better in relative numbers) than those of the second best method, Unity fingerprints and binary kernel discrimination. In 10 out of 11 sets of active compounds the combination of atom environments and the naive Bayesian classifier appears to be the superior method, while in the remaining dataset, data fusion and binary kernel discrimination in combination with Unity fingerprints is the method of choice. Binary kernel discrimination in combination with Unity fingerprints generally comes second in performance overall. The difference in performance can largely be attributed to the different molecular descriptors used. Atom environments outperform Unity fingerprints by a large margin if the combination of these descriptors with the Tanimoto coefficient is compared. The naive Bayesian classifier in combination with information-gain-based feature selection and selection of a sensible number of features performs about as well as binary kernel discrimination in experiments where these classification methods are compared. When used on a monoaminooxidase dataset, atom environments and the naive Bayesian classifier perform as well as binary kernel discrimination in the case of a 50/50 split of training and test compounds. In the case of sparse training data, binary kernel discrimination is found to be superior on this particular dataset. On a third dataset, the atom environment descriptor shows higher retrieval rates than other 2D fingerprints tested here when used in combination with the Tanimoto similarity coefficient. Feature selection is shown to be a crucial step in determining the performance of the algorithm. The representation of molecules by atom environments is found to be more effective than Unity fingerprints for the type of biological receptor similarity calculations examined here. Combining information prior to scoring and including information about inactive compounds, as in the Bayesian classifier and binary kernel discrimination, is found to be superior to posterior data fusion (in the datasets tested here).