Understanding and classifying metabolite space and metabolite-likeness.

Understanding and classifying metabolite space and metabolite-likeness.
复制标题

DOI:
10.1371/journal.pone.0028966
复制
发表时间:
2011
期刊:
影响因子:
3.7
通讯作者:
Hankemeier T
Hankemeier T
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Peironcely JE;Reijmers T;Coulier L;Bender A;Hankemeier T

文献摘要

参考文献

被引文献

相似文献

虽然整个“化学空间”是巨大的(并假设包含1.063到10200个‘小分子’),但这个空间的不同子集仍然可以根据某些结构参数来定义。这种子空间的一个例子是由内源性代谢物跨越的化学空间,内源性代谢物被定义为生物体新陈代谢的‘自然发生’产物。为了更详细地了解这部分化学空间,我们从两个方面分析了人类代谢物占据的化学空间。首先,为了更好地了解代谢物空间,我们对代谢物和非代谢物进行了主成分分析、层次聚类和支架分析,以分析这两类化合物的哪些化学特征是其特征。在这里,我们发现杂原子(氧和氮)的含量以及特定环系的存在能够区分这两类化合物。其次,我们确定了哪些分子描述符和分类器能够区分代谢物和非代谢物,这是通过分配一个类似代谢物的分数来实现的。结果表明,MDL公钥和随机森林相结合的分类效果最好,AUC值为99.13%,特异度为99.84%,选择性为88.79%。这一性能略好于以前的分类器;有趣的是,我们发现药物占据了两个不同的代谢物相似区域,一个是更多的“合成”,另一个是更多的“代谢物相似”。此外,在457个化合物的真实预期数据集上,获得了95.84%的正确分类。总体而言,我们相信,我们为代谢物分类以及更好地了解代谢物化学空间的任务做出了贡献。这一知识现在可用于开发需要类似代谢物的新药,并可用于我们的工作,特别是在代谢组学领域的代谢物鉴定过程中评估候选分子的代谢物相似性。
While the entirety of ‘Chemical Space’ is huge (and assumed to contain between 1063 and 10200 ‘small molecules’), distinct subsets of this space can nonetheless be defined according to certain structural parameters. An example of such a subspace is the chemical space spanned by endogenous metabolites, defined as ‘naturally occurring’ products of an organisms' metabolism. In order to understand this part of chemical space in more detail, we analyzed the chemical space populated by human metabolites in two ways. Firstly, in order to understand metabolite space better, we performed Principal Component Analysis (PCA), hierarchical clustering and scaffold analysis of metabolites and non-metabolites in order to analyze which chemical features are characteristic for both classes of compounds. Here we found that heteroatom (both oxygen and nitrogen) content, as well as the presence of particular ring systems was able to distinguish both groups of compounds. Secondly, we established which molecular descriptors and classifiers are capable of distinguishing metabolites from non-metabolites, by assigning a ‘metabolite-likeness’ score. It was found that the combination of MDL Public Keys and Random Forest exhibited best overall classification performance with an AUC value of 99.13%, a specificity of 99.84% and a selectivity of 88.79%. This performance is slightly better than previous classifiers; and interestingly we found that drugs occupy two distinct areas of metabolite-likeness, the one being more ‘synthetic’ and the other being more ‘metabolite-like’. Also, on a truly prospective dataset of 457 compounds, 95.84% correct classification was achieved. Overall, we are confident that we contributed to the tasks of classifying metabolites, as well as to understanding metabolite chemical space better. This knowledge can now be used in the development of new drugs that need to resemble metabolites, and in our work particularly for assessing the metabolite-likeness of candidate molecules during metabolite identification in the metabolomics field.
DOI: 10.1021/ci010132r
发表时间: 2002-11-01
期刊: JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子: --
作者:
Durant, JL;Leland, BA;Nourse, JG
通讯作者: Nourse, JG
DOI: 10.1016/s0165-9936(03)00607-1
发表时间: 2003-06-01
影响因子: 13.1
作者:
Baumann, K
通讯作者: Baumann, K
DOI: 10.1093/bioinformatics/btr079
发表时间: 2011-04-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Brown M;Wedge DC;Goodacre R;Kell DB;Baker PN;Kenny LC;Mamas MA;Neyses L;Dunn WB
通讯作者: Dunn WB
DOI: 10.1021/ci700286x
发表时间: 2008-01-01
影响因子: 5.6
作者:
Ertl, Peter;Roggo, Silvio;Schuffenhauer, Ansgar
通讯作者: Schuffenhauer, Ansgar
DOI: 10.1016/s0169-7439(00)00056-3
发表时间: 2000-05-08
影响因子: 3.9
作者:
Badertscher, M;Korytko, A;Pretsch, E
通讯作者: Pretsch, E