How large is the metabolome? A critical analysis of data exchange practices in chemistry.

How large is the metabolome? A critical analysis of data exchange practices in chemistry.
复制标题

DOI:
10.1371/journal.pone.0005440
复制
发表时间:
2009
期刊:
影响因子:
3.7
通讯作者:
Fiehn O
Fiehn O
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Kind T;Scholz M;Fiehn O

文献摘要

参考文献

被引文献

相似文献

通过基因组引导的代谢途径重建来计算物种的代谢组大小会错过来自孤儿基因和来自缺乏注释基因的酶的所有产物。因此,代谢组需要通过实验确定。如果可以查询经过同行评审的公共数据库,以汇编已报告的给定物种结构的目标列表,则质谱法的注释将大大受益。我们详细介绍了目前的障碍,编译这样一个知识基础的代谢产物。作为一个例子,水稻的结果。水稻的两个亚种,日本稻和印度稻已经完全测序。比较了几个主要的小分子数据库,以列出已知的水稻代谢物,包括PubChem、Chemical Abstracts、贝尔斯坦、专利数据库、天然产物词典、SetupX/BinBase、KNApSAcK DB,以及最后通过计算方法获得的那些数据库,即RiceCyc、KEGG和Reactome。在搜索这些数据库时,检索到了5,000多个小分子。不幸的是,大多数情况下,真正的水稻代谢物与非代谢物数据库条目(如农药)一起检索。数据库化合物列表的重叠很难比较,因为结构不是以机器可读格式编码的,或者因为化合物标识符在数据库之间没有交叉引用。我们的结论是,目前的数据库不能全面检索所有已知的代谢产物。代谢组列表大多局限于基因组重建途径。我们建议(生物)化学数据库的提供商将其数据库标识符丰富为PubChem ID和InChIKeys,以实现跨数据库查询。此外,同行评审的期刊库需要强制提交机器可读格式的结构和光谱,以允许对包含化学结构的文章进行自动语义注释。出版标准和数据库架构的这种变化将使研究人员能够汇编有关物种代谢组的现有知识,这些知识可能会扩展到衍生信息,如光谱库,器官特异性代谢物和交叉研究比较。
Calculating the metabolome size of species by genome-guided reconstruction of metabolic pathways misses all products from orphan genes and from enzymes lacking annotated genes. Hence, metabolomes need to be determined experimentally. Annotations by mass spectrometry would greatly benefit if peer-reviewed public databases could be queried to compile target lists of structures that already have been reported for a given species. We detail current obstacles to compile such a knowledge base of metabolites. As an example, results are presented for rice. Two rice (oryza sativa) subspecies have been fully sequenced, oryza japonica and oryza indica. Several major small molecule databases were compared for listing known rice metabolites comprising PubChem, Chemical Abstracts, Beilstein, Patent databases, Dictionary of Natural Products, SetupX/BinBase, KNApSAcK DB, and finally those databases which were obtained by computational approaches, i.e. RiceCyc, KEGG, and Reactome. More than 5,000 small molecules were retrieved when searching these databases. Unfortunately, most often, genuine rice metabolites were retrieved together with non-metabolite database entries such as pesticides. Overlaps from database compound lists were very difficult to compare because structures were either not encoded in machine-readable format or because compound identifiers were not cross-referenced between databases. We conclude that present databases are not capable of comprehensively retrieving all known metabolites. Metabolome lists are yet mostly restricted to genome-reconstructed pathways. We suggest that providers of (bio)chemical databases enrich their database identifiers to PubChem IDs and InChIKeys to enable cross-database queries. In addition, peer-reviewed journal repositories need to mandate submission of structures and spectra in machine readable format to allow automated semantic annotation of articles containing chemical structures. Such changes in publication standards and database architectures will enable researchers to compile current knowledge about the metabolome of species, which may extend to derived information such as spectral libraries, organ-specific metabolites, and cross-study comparisons.
DOI: 10.1101/gr.1212003
发表时间: 2003-11-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Arita, M
通讯作者: Arita, M
DOI: 10.1021/ci60024a001
发表时间: 1980-01-01
期刊: JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子: --
作者:
BAKER, DB;HORISZNY, JW;METANOMSKI, WV
通讯作者: METANOMSKI, WV
DOI: 10.1021/ci060139c
发表时间: 2006-11-27
影响因子: 5.6
作者:
Casher, Omer;Rzepa, Henry S.
通讯作者: Rzepa, Henry S.
DOI: 10.1021/ci050400b
发表时间: 2006-05
影响因子: 5.6
作者:
Guha R;Howard MT;Hutchison GR;Murray-Rust P;Rzepa H;Steinbeck C;Wegner J;Willighagen EL
通讯作者: Willighagen EL
DOI: 10.1021/ci00013a010
发表时间: 1993-05-01
期刊: JOURNAL OF CHEMICAL INFORMATION AND COMPUTER SCIENCES
影响因子: --
作者:
IBISON, P;JACQUOT, M;JOHNSON, AP
通讯作者: JOHNSON, AP