Retip: Retention Time Prediction for Compound Annotation in Untargeted Metabolomics.

Retip: Retention Time Prediction for Compound Annotation in Untargeted Metabolomics.
复制标题

DOI:
10.1021/acs.analchem.9b05765
复制
发表时间:
2020-06-02
影响因子:
7.4
通讯作者:
Fiehn O
Fiehn O
中科院分区:
化学1区
文献类型:
--
作者:
Bonini P;Kind T;Tsugawa H;Barupal DK;Fiehn O

文献摘要

参考文献

被引文献

相似文献

在LC-MS/MS的非靶向代谢组学中,未确定的峰仍然是一个主要问题,通过结合MS/MS匹配和保留时间,峰注释的可信度增加。我们在这里展示了如何根据分子结构来预测保留时间。两个大型的公开可用的数据集用于机器学习的模型训练:981个初级代谢物和生物胺的Fiehn亲水相互作用液相色谱数据集(HILIC),以及使用反相液相色谱(RPLC)的852个次级代谢物的RIKEN植物专门代谢组注释(PLUMP)数据库。ReTip R包中集成了五种不同的机器学习算法:随机森林、贝叶斯正则化神经网络、XGBoost、光梯度助推机(LightGBM)和用于建立保留时间预测模型的Kera算法。在R中开发了一个完整的保留时间预测工作流,可以从GitHub存储库(https://www.retip.app).)免费下载KERAS在测试集中的表现优于其他机器学习算法,具有最小的过拟合,通过训练、测试和验证集之间的微小误差差异得到验证。KERAS的平均绝对误差HILIC为0.78分钟,RPLC为0.57分钟。RETIP被集成到质谱学软件工具MS-Dial和MS-Finder中,从而实现完整的复合注释工作流程。在对小鼠血浆样本的测试应用中,我们发现在MS-finder化合物鉴定软件中搜索所有异构体时,候选结构的数量减少了68%。保留时间预测提高了液相色谱分析的识别率,从而改进了代谢组学数据的生物学解释。
Unidentified peaks remain a major problem in untargeted metabolomics by LC-MS/MS. Confidence in peak annotations increases by combining MS/MS matching and retention time. We here show how retention times can be predicted from molecular structures. Two large, publicly available data sets were used for model training in machine learning: the Fiehn hydrophilic interaction liquid chromatography data set (HILIC) of 981 primary metabolites and biogenic amines, and the RIKEN plant specialized metabolome annotation (PlaSMA) database of 852 secondary metabolites that uses reversed-phase liquid chromatography (RPLC). Five different machine learning algorithms have been integrated into the Retip R package: the random forest, Bayesian-regularized neural network, XGBoost, light gradient-boosting machine (LightGBM), and Keras algorithms for building the retention time prediction models. A complete workflow for retention time prediction was developed in R. It can be freely downloaded from the GitHub repository (https://www.retip.app). Keras outperformed other machine learning algorithms in the test set with minimum overfitting, verified by small error differences between training, test, and validation sets. Keras yielded a mean absolute error of 0.78 min for HILIC and 0.57 min for RPLC. Retip is integrated into the mass spectrometry software tools MS-DIAL and MS-FINDER, allowing a complete compound annotation workflow. In a test application on mouse blood plasma samples, we found a 68% reduction in the number of candidate structures when searching all isomers in MS-FINDER compound identification software. Retention time prediction increases the identification rate in liquid chromatography and subsequently leads to an improved biological interpretation of metabolomics data.
DOI: 10.1021/acs.analchem.6b00770
发表时间: 2016-08-16
影响因子: 7.4
作者:
Tsugawa H;Kind T;Nakabayashi R;Yukihira D;Tanaka W;Cajka T;Saito K;Fiehn O;Arita M
通讯作者: Arita M
DOI: 10.1021/acs.analchem.8b01527
发表时间: 2018-09-18
影响因子: 7.4
作者:
Blazenovic, Ivana;Shen, Tong;Fiehn, Oliver
通讯作者: Fiehn, Oliver
DOI: 10.1021/acs.analchem.6b02075
发表时间: 2016-10-04
影响因子: 7.4
作者:
Falchi, Federico;Bertozzi, Sine Mandrup;Armirotti, Andrea
通讯作者: Armirotti, Andrea
DOI: 10.1021/acs.analchem.8b03118
发表时间: 2018-11-06
影响因子: 7.4
作者:
Samaraweera MA;Hall LM;Hill DW;Grant DF
通讯作者: Grant DF
DOI: 10.1016/j.phytochem.2014.10.005
发表时间: 2014-12-01
期刊: PHYTOCHEMISTRY
影响因子: 3.8
作者:
Eugster, Philippe J.;Boccard, Julien;Carrupt, Pierre-Alain
通讯作者: Carrupt, Pierre-Alain