Reoptimization of MDL keys for use in drug discovery

Reoptimization of MDL keys for use in drug discovery
复制标题

DOI:
10.1021/ci010132r
复制
发表时间:
2002-11-01
期刊:
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子:
--
通讯作者:
Nourse, JG
Nourse, JG
中科院分区:
其他
文献类型:
--
作者:
Durant, JL;Leland, BA;Nourse, JG

文献摘要

被引文献

相似文献

多年来,MDL产品已经公开了基于2D描述符的166位和960位密钥集。这些密钥集最初是为子结构搜索而构建和优化的。我们报告的MDL键集的性能进行了重新优化,用于分子相似性的改进。957种化合物的测试数据集的分类性能从166位密钥集的0.65和960位密钥集的0.67增加到包含208位的序列S/N修剪密钥集的0.71和包含548位的遗传算法优化密钥集的0.71。我们提出了一个概述的底层技术支持描述符的定义和编码这些描述符到键集。该技术允许将描述符定义为各种拓扑分离处的原子性质、键性质和原子邻域的组合,以及支持许多自定义描述符。然后,这些描述符可以用于设置密钥集中的一个或多个位。我们构建了各种密钥集,并优化了它们在生物活性物质聚类中的性能。使用Briem和Lessel开发的方法测量性能。“定向修剪”是通过基于随机选择、比特的对数值或比特的对数S/N比值从密钥集中消除比特来执行的。随机修剪实验强调了密钥集长度超过1000位时密钥集性能的不敏感性。与最初的预期相反,基于各个比特的重复值的修剪导致密钥集的性能低于随机修剪产生的密钥集。相比之下,修剪的基础上的prossal SIN比率被发现产生的密钥集,表现优于那些从随机修剪。我们还探讨了使用遗传算法在选择最佳的密钥集。再一次,性能只是键集大小的弱函数,并且优化未能识别出单个全局最优键集。相反,可以产生多个同样最优的密钥集,这些密钥集具有它们编码的描述符的相对低的重叠。
For a number of years MDL products have exposed both 166 bit and 960 bit keysets based on 2D descriptors. These keysets were originally constructed and optimized for substructure searching. We report on improvements in the performance of MDL keysets which are reoptimized for use in molecular similarity. Classification performance for a test data set of 957 compounds was increased from 0.65 for the 166 bit keyset and 0.67 for the 960 bit keyset to 0.71 for a surprisal S/N pruned keyset containing 208 bits and 0.71 for a genetic algorithm optimized keyset containing 548 bits. We present an overview of the underlying technology supporting the definition of descriptors and the encoding of these descriptors into keysets. This technology allows definition of descriptors as combinations of atom properties, bond properties, and atomic neighborhoods at various topological separations as well as supporting a number of custom descriptors. These descriptors can then be used to set one or more bits in a keyset. We constructed various keysets and optimized their performance in clustering bioactive substances. Performance was measured using methodology developed by Briem and Lessel. "Directed pruning" was carried out by eliminating bits from the keysets on the basis of random selection, values of the surprisal of the bit, or values of the surprisal S/N ratio of the bit. The random pruning experiment highlighted the insensitivity of keyset performance for keyset lengths of more than 1000 bits. Contrary to initial expectations, pruning on the basis of the surprisal values of the various bits resulted in keysets which underperformed those resulting from random pruning. In contrast, pruning on the basis of the surprisal SIN ratio was found to yield keysets which performed better than those resulting from random pruning. We also explored the use of genetic algorithms in the selection of optimal keysets. Once more the performance was only a weak function of keyset size, and the optimizations failed to identify a single globally optimal keyset. Instead multiple, equally optimal keysets could be produced which had relatively low overlap of the descriptors they encoded.