Comparison of Nonbinary Similarity Coefficients for Similarity Searching, Clustering and Compound Selection

Comparison of Nonbinary Similarity Coefficients for Similarity Searching, Clustering and Compound Selection
复制标题

DOI:
10.1021/ci8004644
复制
发表时间:
2009-05-01
影响因子:
5.6
通讯作者:
Holliday, John
Holliday, John
中科院分区:
化学2区
文献类型:
--
作者:
Al Khalifa, Aysha;Haranczyk, Maciej;Holliday, John

文献摘要

被引文献

相似文献

最近的几项研究比较了选择的相似性系数时,适用于二进制指纹表示的化学数据库的相对性能。当用于基于相似性的技术(如相似性搜索、数据库聚类和基于相异性的化合物选择)时,性能的相当大的变化已经被报道,其原因与分子大小密切相关。对于许多这些相似性系数,可以推导出一种替代形式,它适用于非二进制数据集,如计算或测量的物理化学性质,或亚结构片段的计数。在这里,我们报告的几项研究已进行调查的相对性能的12个系数时,应用到非二进制数据使用这种(DIS)相似性为基础的技术。结果表明,没有一个单一的系数是适合所有的方法研究和二进制数据检测的大小偏差是不明显的数据时,因此,系数是非二进制的性质。
Several recent studies have compared the relative performance of a selection of similarity coefficients when applied to chemical databases represented by binary fingerprints. Considerable variation in performance, when used for (dis)similarity-based techniques, such as similarity searching, database clustering, and dissimilarity-based compound selection, has been reported, the reasons for which are closely related to molecular size. For many of these similarity coefficients, an alternative form can be derived which is applicable to sets of nonbinary data, such as calculated or measured physicochemical properties, or counts of substructural fragments. Here we report on several studies which have been undertaken to investigate the relative performance of twelve coefficients when applied to nonbinary data using such (dis)similarity-based techniques. Results suggest that no single coefficient is appropriate for all methodologies investigated and that the size bias detected with binary data is not as apparent when the data and, hence, coefficient are nonbinary in nature.