A constant time algorithm for estimating the diversity of large chemical libraries

A constant time algorithm for estimating the diversity of large chemical libraries
复制标题

DOI:
10.1021/ci000091j
复制
发表时间:
2001-01-01
期刊:
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子:
--
通讯作者:
Agrafiotis, DK
Agrafiotis, DK
中科院分区:
其他
文献类型:
--
作者:
Agrafiotis, DK

文献摘要

被引文献

相似文献

我们描述了一种新颖的多样性度量,用于组合化学和高通量筛选实验的设计。该方法估计感兴趣集合中分子间差异的累积概率分布,然后使用 Kolmogorov-Smirnov 统计量测量该分布与均匀样本的相应分布的偏差。这种方法的显着优点是可以使用概率采样轻松估计累积分布,并且不需要详尽枚举数据集中的所有成对距离。该函数直观,计算速度非常快,不依赖于集合的大小,并且可用于在全局和局部尺度上执行多样性估计。更重要的是,它允许对不同基数的数据集进行有意义的比较,并且不受维数灾难的影响,维数灾难困扰着许多其他多样性指数。使用组合化学文献中的示例证明了这种方法的优点。
We describe a novel diversity metric for use in the design of combinatorial chemistry and high-throughput screening experiments. The method estimates the cumulative probability distribution of intermolecular dissimilarities in the collection of interest and then measures the deviation of that distribution from the respective distribution of a uniform sample using the Kolmogorov-Smirnov statistic. The distinct advantage of this approach is that the cumulative distribution can be easily estimated using probability sampling and does not require exhaustive enumeration of all pairwise distances in the data set. The function is intuitive, very fast to compute, does not depend on the size of the collection, and can be used to perform diversity estimates on both global and local scale. More importantly, it allows meaningful comparison of data sets of different cardinality and is not affected by the curse of dimensionality, which plagues many other diversity indices. The advantages of this approach are demonstrated using examples from the combinatorial chemistry literature.