Towards de novo identification of metabolites by analyzing tandem mass spectra

Towards de novo identification of metabolites by analyzing tandem mass spectra
复制标题

DOI:
10.1093/bioinformatics/btn270
复制
发表时间:
2008-08-15
期刊:
影响因子:
5.8
通讯作者:
Rasche, Florian
Rasche, Florian
中科院分区:
生物学3区
文献类型:
--
作者:
Boecker, Sebastian;Rasche, Florian

文献摘要

被引文献

相似文献

动机:质谱是蛋白质组学和代谢组学中最广泛使用的技术之一。作为一种高通量方法,它产生了大量数据,需要对光谱进行自动分析。显然,可以轻松地采用用于蛋白质分析的数据库搜索方法来分析代谢物质谱。但是对于代谢产物而言,从头开始的解释比对于蛋白质数据更重要,因为代谢物光谱数据库仅涵盖一小部分天然发生的代谢物:甚至模型植物拟南芥具有大量的酶,其底物和产品仍然未知。生物培训领域搜索可能用作药物的代谢物的生物学不同领域。从头开始鉴定代谢物质谱需要新的概念和方法,因为与蛋白质不同,代谢物具有非线性分子结构。反应:在这项工作中,我们介绍了一种从tandem质谱的新自动化的DE鉴定的方法。通常认为质谱数据不足以鉴定分​​子结构,因此我们要估计未知代谢物的分子公式,这是其鉴定的关键步骤。该方法首先计算所有解释母峰质量的分子公式。然后,在碎片化质谱​​中对应于所有峰的分子公式对应于分子公式,而边缘对应于假设的碎片步骤。之后,我们的算法计算了此图的最大评分子树:光谱中的每个峰必须最多得分一次,因此子树应仅包含一个每个峰的解释。不幸的是,找到此子树是NP-HARD。我们建议三种精确的算法(包括一种固定参数可拖动算法)以及两个启发式方法来解决该问题。对实际质谱的测试表明,FPT算法和启发式方法可以快速解决该问题,并提供了出色的结果:对于所有32种测试化合物,正确的解决方案都是前五名的建议,对于26种化合物,确切算法的第一个建议是正确的。 。
Motivation: Mass spectrometry is among the most widely used technologies in proteomics and metabolomics. Being a high-throughput method, it produces large amounts of data that necessitates an automated analysis of the spectra. Clearly, database search methods for protein analysis can easily be adopted to analyze metabolite mass spectra. But for metabolites, de novo interpretation of spectra is even more important than for protein data, because metabolite spectra databases cover only a small fraction of naturally occurring metabolites: even the model plant Arabidopsis thaliana has a large number of enzymes whose substrates and products remain unknown. The field of bio-prospection searches biologically diverse areas for metabolites which might serve as pharmaceuticals. De novo identification of metabolite mass spectra requires new concepts and methods since, unlike proteins, metabolites possess a non-linear molecular structure.Results: In this work, we introduce a method for fully automated de novo identification of metabolites from tandem mass spectra. Mass spectrometry data is usually assumed to be insufficient for identification of molecular structures, so we want to estimate the molecular formula of the unknown metabolite, a crucial step for its identification. The method first calculates all molecular formulas that explain the parent peak mass. Then, a graph is build where vertices correspond to molecular formulas of all peaks in the fragmentation mass spectra, whereas edges correspond to hypothetical fragmentation steps. Our algorithm afterwards calculates the maximum scoring subtree of this graph: each peak in the spectra must be scored at most once, so the subtree shall contain only one explanation per peak. Unfortunately, finding this subtree is NP-hard. We suggest three exact algorithms (including one fixed-parameter tractable algorithm) as well as two heuristics to solve the problem. Tests on real mass spectra show that the FPT algorithm and the heuristics solve the problem suitably fast and provide excellent results: for all 32 test compounds the correct solution was among the top five suggestions, for 26 compounds the first suggestion of the exact algorithm was correct.