How Similar Are Similarity Searching Methods? A Principal Component Analysis of Molecular Descriptor Space

How Similar Are Similarity Searching Methods? A Principal Component Analysis of Molecular Descriptor Space
复制标题

DOI:
10.1021/ci800249s
复制
发表时间:
2009-01-01
影响因子:
5.6
通讯作者:
Davies, John W.
Davies, John W.
中科院分区:
化学2区
文献类型:
--
作者:
Bender, Andreas;Jenkins, Jeremy L.;Davies, John W.

文献摘要

被引文献

相似文献

不同的分子描述符捕捉分子结构的不同方面,但这种影响尚未被大规模系统地量化。在这项工作中,我们通过重复选择查询化合物并对数据库的其余部分进行排名来计算37个描述符的相似性。计算不同描述符的等级排序之间的欧几里得距离以确定描述符(与复合描述符相对)相似性,然后进行PCA以进行可视化。四个广泛的描述符类,这是循环指纹;循环指纹考虑计数;基于路径和键控指纹;和pharmacophoric描述符。描述符行为更多地由这四个类定义,而不是特定的参数化。使用计数而不是指纹的存在/不存在显着改变描述符的行为,这是至关重要的拓扑自相关向量的性能,但不是循环指纹。四点药效团(piDAPH 4)令人惊讶地导致比三点药效团高得多的检索率(28.21%对19.15%),但仍然具有相似的化合物排序(相似活性物质的检索)。如果存在复杂的环系统或分支模式,则观察个体排名,环状指纹似乎比基于路径的指纹更合适;基于计数的指纹可能更适合于具有大量重复亚基(酰胺键、糖环、萜烯)的数据库。基于信息选择不同指纹进行一致性评分(ECFP 4/TGD指纹)仅导致单一指纹结果的边际改善。虽然利用正交描述符行为来提高一致性虚拟筛选中的检索率似乎是不平凡的,但这些描述符仍然各自检索不同的活性物质,这证实了在前瞻性虚拟筛选设置中单独采用不同描述符的策略。
Different molecular descriptors capture different aspects of molecular structures, but this effect has not yet been quantified systematically on a large scale. In this work, we calculate the similarity of 37 descriptors by repeatedly selecting query compounds and ranking the rest of the database. Euclidean distances between the rank-ordering of different descriptors are calculated to determine descriptor (as opposed to compound) similarity, followed by PCA for visualization. Four broad descriptor classes are identified, which are circular fingerprints; circular fingerprints considering counts; path-based and keyed fingerprints; and pharmacophoric descriptors. Descriptor behavior is much more defined by those four classes than the particular parametrization. Using counts instead of the presence/absence of fingerprints significantly changes descriptor behavior, which is crucial for performance of topological autocorrelation vectors, but not circular fingerprints. Four-point pharmacophores (piDAPH4) surprisingly lead to much higher retrieval rates than three-point pharmacophores (28.21% vs 19.15%) but still similar rank-ordering of compounds (retrieval of similar actives). Looking into individual rankings, circular fingerprints seem more appropriate than path-based fingerprints if complex ring systems or branching patterns are present; count-based fingerprints could be more suitable in databases with a large number of repeated subunits (amide bonds, sugar rings, terpenes). Information-based selection of diverse fingerprints for consensus scoring (ECFP4/TGD fingerprints) led only to marginal improvement over single fingerprint results. While it seems to be nontrivial to exploit orthogonal descriptor behavior to improve retrieval rates in consensus virtual screening, those descriptors still each retrieve different actives which corroborates the strategy of employing diverse descriptors individually in prospective virtual screening settings.