Finding Correspondence between Metabolomic Features in Untargeted Liquid Chromatography-Mass Spectrometry Metabolomics Datasets.

Finding Correspondence between Metabolomic Features in Untargeted Liquid Chromatography-Mass Spectrometry Metabolomics Datasets.
复制标题

DOI:
10.1021/acs.analchem.1c03592
复制
发表时间:
2022-04-12
影响因子:
7.4
通讯作者:
Ebbels, Timothy
Ebbels, Timothy
中科院分区:
化学1区
文献类型:
--
作者:
Pinto, Rui Climaco;Karaman, Ibrahim;Lewis, Matthew R.;Hallqvist, Jenny;Kaluarachchi, Manuja;Graca, Goncalo;Chekmeneva, Elena;Durainayagam, Brenan;Ghanbari, Mohsen;Ikram, M. Arfan;Zetterberg, Henrik;Griffin, Julian;Elliott, Paul;Tzoulaki, Ioanna;Dehghan, Abbas;Herrington, David;Ebbels, Timothy

文献摘要

参考文献

相似文献

多个数据集的整合可以极大地增强生物分析研究,例如,通过增加发现和验证生物标志物的能力。在液相色谱-质谱(LC-MS)代谢组学中,由于大多数代谢组学特征没有注释,因此无法通过化学身份进行匹配,因此特别难以联合收割机组合非目标数据集。通常,每个特征可用的信息是保留时间(RT)、质荷比(m/z)和特征强度(FI)。来自不同数据集中相同代谢物的成对特征可能表现出微小但显著的差异,这使得匹配非常具有挑战性。目前解决这一问题的方法过于简单或依赖于无法在所有情况下满足的假设。我们提出了一种方法来找到两个相似的LC-MS代谢组学实验或批次之间的特征对应关系,仅使用特征的RT,m/z和FI。我们证明了真实的和合成数据集的方法,使用六个正交验证策略来衡量匹配质量。在我们的主要示例中,4953个特征是唯一匹配的,其中604个手动注释的特征中有585个(96.8%)是正确的。在第二个例子中,2324个特征可以唯一匹配,87个注释特征中有79个(90.8%)正确匹配。大多数错过的注释匹配是在行为与RT,MZ和FI的建模数据集间偏移非常不同的特征之间。在第三个示例中,每个数据集具有4755个特征的模拟数据,99.6%的匹配是正确的。最后,使用我们的方法匹配其他三个数据集对的结果进行了比较,与已发表的替代方法,metabCombiner,显示我们的方法的优势。该方法可以使用M2 S(Match 2 Sets)应用,这是一个免费的开源MATLAB工具箱,可在
Integration of multiple datasets can greatly enhance bioanalytical studies, for example, by increasing power to discover and validate biomarkers. In liquid chromatography–mass spectrometry (LC–MS) metabolomics, it is especially hard to combine untargeted datasets since the majority of metabolomic features are not annotated and thus cannot be matched by chemical identity. Typically, the information available for each feature is retention time (RT), mass-to-charge ratio (m/z), and feature intensity (FI). Pairs of features from the same metabolite in separate datasets can exhibit small but significant differences, making matching very challenging. Current methods to address this issue are too simple or rely on assumptions that cannot be met in all cases. We present a method to find feature correspondence between two similar LC–MS metabolomics experiments or batches using only the features’ RT, m/z, and FI. We demonstrate the method on both real and synthetic datasets, using six orthogonal validation strategies to gauge the matching quality. In our main example, 4953 features were uniquely matched, of which 585 (96.8%) of 604 manually annotated features were correct. In a second example, 2324 features could be uniquely matched, with 79 (90.8%) out of 87 annotated features correctly matched. Most of the missed annotated matches are between features that behave very differently from modeled inter-dataset shifts of RT, MZ, and FI. In a third example with simulated data with 4755 features per dataset, 99.6% of the matches were correct. Finally, the results of matching three other dataset pairs using our method are compared with a published alternative method, metabCombiner, showing the advantages of our approach. The method can be applied using M2S (Match 2 Sets), a free, open-source MATLAB toolbox, available at .
DOI: 10.1007/s11306-015-0893-5
发表时间: 2016-01-01
期刊: METABOLOMICS
影响因子: 3.6
作者:
Ganna, Andrea;Fall, Tove;Ingelsson, Erik
通讯作者: Ingelsson, Erik
DOI: 10.1093/aje/kwf113
发表时间: 2002-11-01
影响因子: 5
作者:
Bild, DE;Bluemke, DA;Tracy, RP
通讯作者: Tracy, RP
DOI: 10.1093/bioinformatics/btaa037
发表时间: 2020-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Wu, Chiung-Ting;Wang, Yizhi;Yu, Guoqiang
通讯作者: Yu, Guoqiang
DOI: 10.1002/jms.1777
发表时间: 2010-07-01
影响因子: 2.3
作者:
Horai, Hisayuki;Arita, Masanori;Nishioka, Takaaki
通讯作者: Nishioka, Takaaki
DOI: 10.1038/nature06882
发表时间: 2008-05-15
期刊: NATURE
影响因子: 64.8
作者:
Holmes, Elaine;Loo, Ruey Leng;Elliott, Paul
通讯作者: Elliott, Paul