Method for Assessing the Statistical Significance of Mass Spectral Similarities Using Basic Local Alignment Search Tool Statistics

Method for Assessing the Statistical Significance of Mass Spectral Similarities Using Basic Local Alignment Search Tool Statistics
复制标题

DOI:
10.1021/ac401564v
复制
发表时间:
2013-09-03
影响因子:
7.4
通讯作者:
Fukusaki, Eiichiro
Fukusaki, Eiichiro
中科院分区:
化学1区
文献类型:
--
作者:
Matsuda, Fumio;Tsugawa, Hiroshi;Fukusaki, Eiichiro

文献摘要

被引文献

相似文献

提出了一种利用改进的基本局部比对搜索工具(BLAST; Karlin- Altschul)统计量评估质谱相似性统计显著性的新方法。在基于气相色谱/质谱的代谢组学中,原始代谢组数据中的许多信号是基于质谱和标准谱之间意想不到的相似性来识别的。由于在所观察到的光谱中不可避免地存在噪声,因此确定的代谢物列表中包含一些假阳性。在电子电离质谱法(blast)中,采用一般评分方案计算两个质谱的相似度评分,并由此计算获得该评分的偶然概率(P值)。为此,提出了将单位El质谱转换为质谱序列的简单规则,并给出了对准质谱序列的分数矩阵。使用随机生成的质谱序列进行蒙特卡罗模拟表明,零分布或期望命中数(E值)遵循改进的Karlin-Altschul统计。用该方法对绿茶提取物的代谢物数据集进行了分析。在代谢组数据中的171个代谢物信号中,有93个信号与参考数据具有显著相似性(P < 0.015)。由于预期假阳性数为2.6,因此估计假发现率为2.8%,表明搜索阈值(P < 0.015)对于代谢物鉴定是合理的。
A novel method for assessing the statistical significance of mass spectral similarities was developed using modified basic local alignment search tool (BLAST; Karlin- Altschul) statistics. In gas chromatography/mass spectrometry-based metabolomics, many signals in raw metabolome data are identified on the basis of unexpected similarities among mass spectra and the spectra of standards. Since there is inevitably noise in the observed spectra, a list of identified metabolites includes some false positives. In the developed method, electron ionization (El) mass spectrometry-BLAST, a similarity score of two mass spectra is calculated using a general scoring scheme, from which the probability of obtaining the score by chance (P value) is calculated. For this purpose, a simple rule for converting a unit El mass spectrum to a mass spectral sequence as well as a score matrix for aligned mass spectral sequences was developed. A Monte Carlo simulation using randomly generated mass spectral sequences demonstrated that the null distribution or the expected number of hits (E value) follows modified Karlin-Altschul statistics. A metabolite data set obtained from green tea extract was analyzed using the developed method. Among 171 metabolite signals in the metabolome data, 93 signals were identified on the basis of significant similarities (P < 0.015) with reference data. Since the expected number of false positives is 2.6, the false discovery rate was estimated to be 2.8%, indicating that the search threshold (P < 0.015) is reasonable for metabolite identification.