Designing libraries with CNS activity

Designing libraries with CNS activity
复制标题

DOI:
10.1021/jm990017w
复制
发表时间:
1999-12-02
影响因子:
7.3
通讯作者:
Murcko, MA
Murcko, MA
中科院分区:
医学1区
文献类型:
--
作者:
Ajay;Bemis, GW;Murcko, MA

文献摘要

被引文献

相似文献

图书馆设计是一项重要而艰巨的任务。在本文中,我们描述了一种可能的解决方案,设计一个CNS主动库。CNS-act;根据是否在数据库中被描述为具有某种CNS活性,从CMC和MDDR数据库中选择伊韦斯和非活性物。该分类方案导致超过15000种活性物和超过50000种非活性物。每个分子由7个ID描述符(分子量、供体数目、受体数目等)描述。和166个2D描述符(存在/不存在诸如NH 2的官能团)。使用贝叶斯方法训练的神经网络可以使用7个1D描述符正确预测约75%的活性和65%的非活性。性能提高到83%的活动集上的预测准确度和79%的非活动上添加2D描述符。在具有275种化合物的数据库中,其中每种化合物的CNS活性是已知的(来自文献),我们分别在活性物和非活性物上实现了92%和71%的准确性。因此,我们构建的模型可以用作“过滤器”来检查化学库中任何一组建议的分子。作为我们方法实用性的一个例子,我们描述了一个小型潜在的中枢神经系统活性分子库的生成,这些分子将适合组合化学。这是通过建立和分析一个由药物分子中常见的框架和侧链构成的一百万种化合物的大型数据库来完成的。
Library design is an important and difficult task. In this paper we describe one possible solution to designing a CNS-active library. CNS-act;ives and -inactives were selected from the CMC and the MDDR databases based on whether they were described as having some kind of CNS activity in the databases, This classification scheme results in over 15 000 actives and over 50 000 inactives. Each molecule is described by 7 ID descriptors (molecular weight, number of donors, number of accepters, etc.) and 166 2D descriptors (presence/absence of functional groups such as NH2). A neural network trained using Bayesian methods can correctly predict about 75% of the actives and 65% of the inactives using the 7 1D descriptors. The performance improves to a prediction accuracy on the active set of 83% and 79% on the inactives on adding the 2D descriptors. On a database with 275 compounds where the CNS activity is known (from the literature) for each compound, we achieve 92% and 71% accuracy on the actives and inactives, respectively. The models we construct can therefore be used as a "filter" to examine any set of proposed molecules in a chemical-library. As an example of the utility of our method, we describe the generation of a small library of potentially CNS-active molecules that would be amenable to combinatorial chemistry. This was done by building and analyzing a large database of a million compounds constructed from frameworks and side chains frequently found in drug molecules.