Consolidating metabolite identifiers to enable contextual and multi-platform metabolomics data analysis.

Consolidating metabolite identifiers to enable contextual and multi-platform metabolomics data analysis.
复制标题

DOI:
10.1186/1471-2105-11-214
复制
发表时间:
2010-04-29
期刊:
影响因子:
3
通讯作者:
Arita M
Arita M
中科院分区:
生物学4区
文献类型:
--
作者:
Redestig H;Kusano M;Fukushima A;Matsuda F;Saito K;Arita M

文献摘要

参考文献

被引文献

相似文献

高通量实验数据的分析取决于描述所测定生物分子的结构良好的数据的可用性。在许多数据分析包中,获得和组织基因、转录本和蛋白质元数据的程序已经简化,但仍然缺乏代谢物元数据。众所周知,化学品标识符是不连贯的,包括范围和覆盖面各不相同的各种不同的参考方案。在线化学品数据库并行使用多种类型的标识符,但缺乏用于可靠数据库整合的通用主键。因此,将实验数据中发现的分析物的标识符与公共数据库中其母体代谢物的标识符相连接可能非常费力。在这里,我们提出了一种策略和一个软件工具,用于整合来自当地参考图书馆和公共数据库的代谢物标识符,这些数据库不依赖于单个共同的主要标识符。该程序构建了分析物和代谢物的互连标识符组,以获得以代谢物为中心的本地SQLite数据库。创建的数据库可用于将内部标识符和同义词映射到外部资源,如KEGG数据库。新的标识符可以导入并直接与现有数据集成。可以从命令行和从统计编程环境R以灵活的方式执行映射,以获得数据集定制的标识符映射。代谢物标识符的高效交叉引用是代谢组学数据分析的关键技术。我们为这项任务提供了一个实用而灵活的解决方案,以及一个开源程序,代谢物掩蔽工具(MetMask),可在www.example.com上获得,它实现了我们的想法。
Analysis of data from high-throughput experiments depends on the availability of well-structured data that describe the assayed biomolecules. Procedures for obtaining and organizing such meta-data on genes, transcripts and proteins have been streamlined in many data analysis packages, but are still lacking for metabolites. Chemical identifiers are notoriously incoherent, encompassing a wide range of different referencing schemes with varying scope and coverage. Online chemical databases use multiple types of identifiers in parallel but lack a common primary key for reliable database consolidation. Connecting identifiers of analytes found in experimental data with the identifiers of their parent metabolites in public databases can therefore be very laborious. Here we present a strategy and a software tool for integrating metabolite identifiers from local reference libraries and public databases that do not depend on a single common primary identifier. The program constructs groups of interconnected identifiers of analytes and metabolites to obtain a local metabolite-centric SQLite database. The created database can be used to map in-house identifiers and synonyms to external resources such as the KEGG database. New identifiers can be imported and directly integrated with existing data. Queries can be performed in a flexible way, both from the command line and from the statistical programming environment R, to obtain data set tailored identifier mappings. Efficient cross-referencing of metabolite identifiers is a key technology for metabolomics data analysis. We provide a practical and flexible solution to this task and an open-source program, the metabolite masking tool (MetMask), available at http://metmask.sourceforge.net, that implements our ideas.
DOI: 10.1186/1471-2105-8-401
发表时间: 2007-10-18
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Cote, Richard G;Jones, Philip;Martens, Lennart;Kerrien, Samuel;Reisinger, Florian;Lin, Quan;Leinonen, Rasko;Apweiler, Rolf;Hermjakob, Henning
通讯作者: Hermjakob, Henning
DOI: 10.1093/nar/gkm882
发表时间: 2008-01
影响因子: 14.9
作者:
Kanehisa M;Araki M;Goto S;Hattori M;Hirakawa M;Itoh M;Katayama T;Kawashima S;Okuda S;Tokimatsu T;Yamanishi Y
通讯作者: Yamanishi Y
DOI: 10.1186/gb-2004-5-10-r80
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Gentleman RC;Carey VJ;Bates DM;Bolstad B;Dettling M;Dudoit S;Ellis B;Gautier L;Ge Y;Gentry J;Hornik K;Hothorn T;Huber W;Iacus S;Irizarry R;Leisch F;Li C;Maechler M;Rossini AJ;Sawitzki G;Smith C;Smyth G;Tierney L;Yang JY;Zhang J
通讯作者: Zhang J
DOI: 10.1093/nar/gkn923
发表时间: 2009-01
影响因子: 14.9
作者:
Huang, Da Wei;Sherman, Brad T.;Lempicki, Richard A.
通讯作者: Lempicki, Richard A.
DOI: 10.1093/nar/gkl838
发表时间: 2007-01
影响因子: 14.9
作者:
Sud M;Fahy E;Cotter D;Brown A;Dennis EA;Glass CK;Merrill AH Jr;Murphy RC;Raetz CR;Russell DW;Subramaniam S
通讯作者: Subramaniam S