In silico enzymatic synthesis of a 400,000 compound biochemical database for nontargeted metabolomics.

In silico enzymatic synthesis of a 400,000 compound biochemical database for nontargeted metabolomics.
复制标题

用于非靶向代谢组学的 400,000 种化合物生化数据库的计算机酶促合成。

DOI:
10.1021/ci400368v
复制
发表时间:
2013
影响因子:
5.6
通讯作者:
Grant,DavidF
Grant,DavidF
中科院分区:
化学2区
文献类型:
--
作者:
Menikarachchi,LochanaC;Hill,DennisW;Hamdalla,MaiA;Mandoiu,IonI;Grant,DavidF

文献摘要

相似文献

目前基于质谱的非靶向代谢组学中的结构鉴定方法依赖于将实验确定的未知化合物的特征与生化数据库中包含的候选化合物的特征相匹配。这种方法的一个主要限制是目前这些数据库中包含的化合物数量相对较少。如果数据库中不存在正确的结构,则无法识别,如果无法识别,则无法将其包含在数据库中。因此,迫切需要使用替代手段用合理设计的生化结构来增强代谢组学数据库。在这里,我们提出了在体内/在硅代谢物数据库(IIMDB),在硅酶促合成的代谢物的数据库,部分解决这个问题。该数据库可在http://metabolomics.pharm.uconn.edu/iimdb/上获得,包括从现有生物化学数据库收集的123000种已知化合物(哺乳动物代谢物、药物、次生植物代谢物和甘油磷脂),以及这些已知化合物的400000多种计算生成的人类I相和II相代谢物。IIMDB具有用户友好的Web界面和程序员友好的RESTful Web服务。IIMDB中95%的计算生成的代谢物在任何现有的数据库中都找不到。然而,21640与PubChem、HMDB、KEGG或HumanCyc中已列出的化合物相同。此外,使用BioSM(一种在化学结构空间中识别生化结构的软件程序)对绝大多数这些计算机模拟代谢物进行生物学评分。这些结果表明,在电脑生化合成代表了一个可行的方法,显着增加非靶向代谢组学应用的生化数据库。
Current methods of structure identification in mass-spectrometry-based nontargeted metabolomics rely on matching experimentally determined features of an unknown compound to those of candidate compounds contained in biochemical databases. A major limitation of this approach is the relatively small number of compounds currently included in these databases. If the correct structure is not present in a database, it cannot be identified, and if it cannot be identified, it cannot be included in a database. Thus, there is an urgent need to augment metabolomics databases with rationally designed biochemical structures using alternative means. Here we present the In Vivo/In Silico Metabolites Database (IIMDB), a database of in silico enzymatically synthesized metabolites, to partially address this problem. The database, which is available at http://metabolomics.pharm.uconn.edu/iimdb/, includes ∼23 000 known compounds (mammalian metabolites, drugs, secondary plant metabolites, and glycerophospholipids) collected from existing biochemical databases plus more than 400 000 computationally generated human phase-I and phase-II metabolites of these known compounds. IIMDB features a user-friendly web interface and a programmer-friendly RESTful web service. Ninety-five percent of the computationally generated metabolites in IIMDB were not found in any existing database. However, 21 640 were identical to compounds already listed in PubChem, HMDB, KEGG, or HumanCyc. Furthermore, the vast majority of these in silico metabolites were scored as biological using BioSM, a software program that identifies biochemical structures in chemical structure space. These results suggest that in silico biochemical synthesis represents a viable approach for significantly augmenting biochemical databases for nontargeted metabolomics applications.