Relationships between molecular complexity, biological activity, and structural diversity

Relationships between molecular complexity, biological activity, and structural diversity
复制标题

DOI:
10.1021/ci0503558
复制
发表时间:
2006-03-01
影响因子:
5.6
通讯作者:
Jacoby, E
Jacoby, E
中科院分区:
化学2区
文献类型:
--
作者:
Schuffenhauer, A;Brown, N;Jacoby, E

文献摘要

被引文献

相似文献

根据Hann等人的理论模型,适度复杂的结构是优选的先导化合物,因为它们导致涉及完整配体分子的特异性结合事件。为了使这一概念在实践中可用于库设计,我们研究了几个复杂性的配体分子的生物活性的措施。我们应用了在Novartis运行的160项试验的历史IC 50/EC 50总结数据,涵盖了各种靶点,其中包括激酶和蛋白酶。GPCR和蛋白质间相互作用。并将其与“无活性”化合物的背景进行比较,所述“无活性”化合物已经筛选了2年,但在任何初步筛选中从未显示出任何活性。作为复杂性的措施,我们使用的结构特征存在于各种分子指纹和描述符的数量。我们发现,随着配体活性的增加,它们的平均复杂性也增加,因此我们可以在每个描述符中建立生物活性所需的最小数量的结构特征。特别适合在这种情况下是Similog密钥和圆形子结构指纹。这些描述符在通过相似性搜索识别生物活性化合物方面也表现得特别好,这表明这些描述符中编码的结构特征与生物活性具有高度相关性。由于特征的数量与分子中存在的原子的数量相关,因此原子的数量也用作合理的复杂性度量,并且较大的分子通常具有较高的活性。由于一方面特征计数和密度与另一方面生物活性之间的关系,存在于几乎所有相似性系数中的尺寸偏差变得特别重要。使用这些系数的多样性选择可以影响所得分子集合的整体复杂性,这对它们表现出的生物活性具有影响。使用基于球体排除的多样性选择方法,如OptiSim与Tanimoto相异度,所得选择的平均特征计数分布向比原始集合更低的复杂度移动,特别是当应用严格的多样性约束时。这种尺寸偏差降低了具有高亚微摩尔活性所需的复杂性的亚组中分子的分数。研究的多样性选择方法,即OptiSim,分裂K均值聚类和自组织映射,都没有产生比随机选择的子集更好的覆盖IC 50汇总数据集的活性空间的子集。
Following the theoretical model by Hann et al. moderately complex structures are preferable lead compounds since they lead to specific binding events involving the complete ligand molecule. To make this concept usable in practice for library design, we studied several complexity measures on the biological activity of ligand molecules. We applied the historical IC50/EC50 summary data of 160 assays run at Novartis covering a diverse range of targets, among them kinases, proteases. GPCRs, and protein-protein interactions. and compared this to the background of "inactive" compounds which have been screened for 2 years but have never shown any activity in any primary screen. As complexity measures we used the number of structural features present in various molecular fingerprints and descriptors. We found generally that with increasing activity of the ligands, their average complexity also increased, and we could therefore establish a minimum number of structural features in each descriptor needed for biological activity. Especially well suited in this context were the Similog keys and circular substructure fingerprints. These are those descriptors, which also perform especially well in the identification of bioactive compounds by similarity search, suggesting that structural features encoded in these descriptors have a high relevance for bioactivity. Since the number of features correlates with the number of atoms present in the molecule, also the number of atoms serves as a reasonable complexity measure and larger molecules have, in general, higher activities. Due to the relationship between feature counts and densities on one hand and biological activity on the other, the size bias present in almost all similarity coefficients becomes especially important. Diversity selections using these coefficients can influence the overall complexity of the resulting set of molecules, which has an impact on the biological activity that they exhibit. Using sphere-exclusion based diversity selection methods, such as OptiSim to-ether with the Tanimoto dissimilarity, the average feature count distribution of the resulting selections is shifted toward lower complexity than that of the original set, particularly when applying tight diversity constraints. This size bias reduces the fraction of molecules in the subsets having the complexity required for a high, sub-micromolar activity. None of the diversity selection methods studied, namely OptiSim, divisive K-means clustering, and self-organizing maps, yielded subsets covering the activity space of the IC50 summary data set better than Subsets selected randomly.