Construction of Confidence Regions for Isotopic Abundance Patterns in LC/MS Data Sets for Rigorous Determination of Molecular Formulas

Construction of Confidence Regions for Isotopic Abundance Patterns in LC/MS Data Sets for Rigorous Determination of Molecular Formulas
复制标题

DOI:
10.1021/ac101278x
复制
发表时间:
2010-09-01
影响因子:
7.4
通讯作者:
Ebbels, Timothy M. D.
Ebbels, Timothy M. D.
中科院分区:
化学1区
文献类型:
--
作者:
Ipsen, Andreas;Want, Elizabeth J.;Ebbels, Timothy M. D.

文献摘要

被引文献

相似文献

人们早就认识到,同位素丰度模式的估计可能有助于确定许多未知的化合物时遇到的进行非靶向代谢分析,使用液相色谱/质谱。虽然已经开发了许多方法来分配启发式得分排名的拟合程度的观察到的丰度模式与理论的,很少的工作已经做了量化的错误,与测量。因此,通常不可能以统计上有意义的方式确定给定的化学式是否可能能够产生观察到的数据。本文提出了一种基于离子到达基本分布构造同位素丰度模式置信域的方法。此外,我们开发了一种方法,这样做,利用汇集在一起的信息从整个色谱峰获得的测量,以及从任何加合物,二聚体,和在质谱中观察到的片段。这大大增加了统计功效,从而使分析人员能够排除潜在的大量候选公式,同时明确防止误报。在实践中,由于探测器饱和和相邻同位素物之间的干扰,可能会与模型假设发生微小偏差。虽然这些因素对统计的严谨性形成了障碍,但它们在很大程度上可以通过将分析限制在中等离子计数和应用稳健的统计方法来克服。使用真实的代谢数据,我们证明了该方法能够减少大量的候选公式的数量,即使没有溴或氯原子存在。我们认为,进一步发展我们的能力,以数学表征的数据可以使更强大的统计分析。
It has long been recognized that estimates of isotopic abundance patterns may be instrumental in identifying the many unknown compounds encountered when conducting untargeted metabolic profiling using liquid chromatography/mass spectrometry. While numerous methods have been developed for assigning heuristic scores to rank the degree of fit of the observed abundance patterns with theoretical ones, little work has been done to quantify the errors that are associated with the measurements made. Thus, it is generally not possible to determine, in a statistically meaningful manner, whether a given chemical formula would likely be capable of producing the observed data. In this paper, we present a method for constructing confidence regions for the isotopic abundance patterns based on the fundamental distribution of the ion arrivals. Moreover, we develop a method for doing so that makes use of the information pooled together from the measurements obtained across an entire chromatographic peak, as well as from any adducts, dimers, and fragments observed in the mass spectra. This greatly increases the statistical power, thus enabling the analyst to rule out a potentially much larger number of candidate formulas while explicitly guarding against false positives. In practice, small departures from the model assumptions are possible due to detector saturation and interferences between adjacent isotopologues. While these factors form impediments to statistical rigor, they can to a large extent be overcome by restricting the analysis to moderate ion counts and by applying robust statistical methods. Using real metabolic data, we demonstrate that the method is capable of reducing the number of candidate formulas by a substantial amount, even when no bromine or chlorine atoms are present. We argue that further developments in our ability to characterize the data mathematically could enable much more powerful statistical analyses.