Mining a chemical database for fragment co-occurrence:: Discovery of "chemical cliches"

Mining a chemical database for fragment co-occurrence:: Discovery of "chemical cliches"
复制标题

DOI:
10.1021/ci050370c
复制
发表时间:
2006-03-01
影响因子:
5.6
通讯作者:
Ijzerman, AP
Ijzerman, AP
中科院分区:
化学2区
文献类型:
--
作者:
Lameijer, EW;Kok, JN;Ijzerman, AP

文献摘要

被引文献

相似文献

如今,已知数百万种不同的化合物,它们的结构存储在电子数据库中。对这些数据的分析可以对化学规律和化学家的习惯产生有价值的见解。因此,我们通过模式搜索搜索了国家癌症研究所(>25万种化合物)的公共数据库。我们将这个数据库的分子拆分成片段,以找出哪些片段存在,它们的频率有多高,以及分子中一个片段的出现是否与另一个不重叠的片段的出现有关。原来,有些碎片和碎片的组合是如此频繁,以至于它们可以被称为“化学陈词滥调”。我们相信,碎片数据可以深入了解合成迄今探索的化学空间。碎片及其(共)赋存状态列表可通过以下方式帮助创建新化合物:(I)系统地列出用于合成新化合物的最受欢迎且因此最容易使用的取代基和环系,(Ii)是适合于铅化合物优化的稀有碎片的易于访问的储存库,以及(Iii)指出化学空间的一些尚未探索的部分。
Nowadays millions of different compounds are known, their structures stored in electronic databases. Analysis of these data could yield valuable insights into the laws of chemistry and the habits of chemists. We have therefore explored the public database of the National Cancer Institute (> 250 000 compounds) by pattern searching. We split the molecules of this database into fragments to find out which fragments exist, how frequent they are, and whether the occurrence of one fragment in a molecule is related to the occurrence of another, nonoverlapping fragment. It turns out that some fragments and combinations of fragments are so frequent that they can be called "chemical cliches". We believe that the fragment data can give insight into the chemical space explored so far by synthesis. The lists of fragments and their (co-)occurrences can help create novel chemical compounds by (i) systematically listing the most popular and therefore most easily used substituents and ring systems for synthesizing new compounds, (ii) being an easily accessible repository for rarer fragments Suitable for lead compound optimization, and (iii) pointing out some of the yet unexplored parts of chemical space.