Unsupervised data base clustering based on Daylight's fingerprint and Tanimoto similarity: A fast and automated way to cluster small and large data sets

Unsupervised data base clustering based on Daylight's fingerprint and Tanimoto similarity: A fast and automated way to cluster small and large data sets
复制标题

DOI:
10.1021/ci9803381
复制
发表时间:
1999-07-01
期刊:
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子:
--
通讯作者:
Butina, D
Butina, D
中科院分区:
其他
文献类型:
--
作者:
Butina, D

文献摘要

被引文献

相似文献

全球制药行业最常用的聚类算法之一是贾维斯-帕特里克(J-P)(Jarvis,R. A. IEEE传输计算1973,C-22,1025-1034)。在Daylight软件下实现J-P,使用Daylight的指纹和Tanimoto相似性指数,可以在几个小时内处理100 k分子的集合。然而,J-P聚类算法有几个相关的问题,这使得它很难在一个一致的和及时的方式聚类大数据集。产生的聚类在很大程度上取决于运行J-P聚类所需的两个参数的选择,使得该方法倾向于产生非常大且异质或同质但太小的聚类。在任何情况下,J-P总是需要耗时的手动调优。本文描述了一种算法,它将识别密集的集群,其中每个集群内的相似性反映了用于聚类的Tanimoto值,更重要的是,其中集群质心将至少相似,在给定的Tanimoto值,集群内的每一个其他分子以一致和自动化的方式。本文中使用的相似性术语反映了两个给定分子之间的总体相似性,如日光指纹和谷本相似性指数所定义的。
One of the most commonly used clustering algorithms within the worldwide pharmaceutical industry is Jarvis-Patrick's (J-P) (Jarvis, R. A. IEEE Trans. Comput. 1973, C-22, 1025-1034). The implementation of J-P under Daylight software, using Daylight's fingerprints and the Tanimoto similarity index, can deal with sets of 100 k molecules in a matter of a few hours. However, the J-P clustering algorithm has several associated problems which make it difficult to cluster large data sets in a consistent and timely manner. The clusters produced are greatly dependent on the choice of the two parameters needed to run J-P clustering, such that this method tends to produce clusters which are either very large and heterogeneous or homogeneous but too small. In any case, J-P always requires time-consuming manual tuning. This paper describes an algorithm which will identify dense clusters where similarity within each cluster reflects the Tanimoto value used for the clustering, and, more importantly, where the cluster centroid will be at least similar, at the given Tanimoto value, to every other molecule within the cluster in a consistent and automated manner. The similarity term used throughout this paper reflects the overall similarity between two given molecules, as defined by Daylight's fingerprints and the Tanimoto similarity index.