The reduced graph descriptor in virtual screening and data-driven clustering of high-throughput screening data

The reduced graph descriptor in virtual screening and data-driven clustering of high-throughput screening data
复制标题

DOI:
10.1021/ci049860f
复制
发表时间:
2004-11-01
期刊:
JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子:
--
通讯作者:
Green, DVS
Green, DVS
中科院分区:
其他
文献类型:
--
作者:
Harper, G;Bravi, GS;Green, DVS

文献摘要

被引文献

相似文献

虚拟筛选和高通量筛选是制药行业中先导化合物发现的两个主要组成部分。在本文中,我们描述了改进以前发表的方法,减少图形的相似性搜索,特别注重基于配体的虚拟筛选,并描述了一种新的使用减少图形的聚类高通量筛选数据。文献中的约简图相似性搜索方法将约简图编码为二进制指纹,这存在许多问题。在本文中,我们扩展了约简图的定义,包括积极和消极的电离组,并介绍了一种新的方法来衡量约简图的相似性的基础上加权编辑距离。超越简单的相似性搜索,我们展示了如何更灵活的查询,可以建立使用简化图和描述一个数据库系统,允许迭代查询与多个表示。简化图捕捉了配体-受体相互作用的许多重要特征,并结合其他全分子描述符,提供了一种信息丰富的方式来审查HTS数据。在这种情况下,我们描述了一种新的使用减少图,介绍了一种方法,我们称之为数据驱动的聚类。其识别由特定的整个分子描述符表示的分子簇,并且富含活性化合物。
Virtual screening and high-throughput screening are two major components of lead discovery within the pharmaceutical industry. In this paper we describe improvements to previously published methods for similarity searching with reduced graphs, with a particular focus on ligand-based virtual screening, and describe a novel use of reduced graphs in the clustering of high-throughput screening data. Literature methods for reduced graph similarity searching encode the reduced graphs as binary fingerprints, which has a number of issues. In this paper we extend the definition of the reduced graph to include positively and negatively ionizable groups and introduce a new method for measuring the similarity of reduced graphs based on a weighted edit distance. Moving beyond simple similarity searching, we show how more flexible queries can be built using reduced graphs and describe a database system that allows iterative querying with multiple representations. Reduced graphs capture many important features of ligand-receptor interactions and, in conjunction with other whole molecule descriptors, provide an informative way to review HTS data. We describe a novel use of reduced graphs in this context, introducing a method we have termed data-driven clustering. that identifies clusters of molecules represented by a particular whole molecule descriptor and enriched in active compounds.