Development of a chemical structure comparison method for integrated analysis of chemical and genomic information in the metabolic pathways

Development of a chemical structure comparison method for integrated analysis of chemical and genomic information in the metabolic pathways
复制标题

DOI:
10.1021/ja036030u
复制
发表时间:
2003-10-01
影响因子:
15
通讯作者:
Kanehisa, M
Kanehisa, M
中科院分区:
化学1区
文献类型:
--
作者:
Hattori, M;Okuno, Y;Kanehisa, M

文献摘要

被引文献

相似文献

细胞功能是由复杂的分子相互作用网络产生的,不仅涉及蛋白质和核酸,还涉及小的化合物。在这里,我们提出了一个有效的算法比较两个化合物的化学结构,其中的化学结构被视为一个图形组成的原子作为节点和共价键作为边缘。基于官能团的概念,针对不同环境的碳、氮、氧等原子种类定义了68种原子种类(节点类型),从而能够检测出具有生物化学意义的特征。两个图的最大公共子图可以通过在关联图中搜索最大团来找到,并且我们引入了算法来加速团的发现和检测最佳局部匹配(简单连接的公共子图)。我们的程序被应用到比较和聚类的9383化合物,主要是代谢化合物,在KEGG/LIGAND数据库。相似化合物的最大簇与碳水化合物相关,并且簇与由KEGG途径图编号表示的途径分类很好地对应。当更详细地检查每个途径图时,可以识别出对应于包含连续反应步骤集的子途径或途径模块的更精细的簇。此外,人们发现,由类似的化合物结构确定的途径模块有时与由基因组背景,即由酶基因的操纵子结构确定的途径模块重叠。
Cellular functions result from intricate networks of molecular interactions, which involve not only proteins and nucleic acids but also small chemical compounds. Here we present an efficient algorithm for comparing two chemical structures of compounds, where the chemical structure is treated as a graph consisting of atoms as nodes and covalent bonds as edges. On the basis of the concept of functional groups, 68 atom types (node types) are defined for carbon, nitrogen, oxygen, and other atomic species with different environments, which has enabled detection of biochemically meaningful features. Maximal common subgraphs of two graphs can be found by searching for maximal cliques in the association graph, and we have introduced heuristics to accelerate the clique finding and to detect optimal local matches (simply connected common subgraphs). Our procedure was applied to the comparison and clustering of 9383 compounds, mostly metabolic compounds, in the KEGG/LIGAND database. The largest clusters of similar compounds were related to carbohydrates, and the clusters corresponded well to the categorization of pathways as represented by the KEGG pathway map numbers. When each pathway map was examined in more detail, finer clusters could be identified corresponding to subpathways or pathway modules containing continuous sets of reaction steps. Furthermore, it was found that the pathway modules identified by similar compound structures sometimes overlap with the pathway modules identified by genomic contexts, namely, by operon structures of enzyme genes.