Large-Scale Prediction of Collision Cross-Section Values for Metabolites in Ion Mobility-Mass Spectrometry

Large-Scale Prediction of Collision Cross-Section Values for Metabolites in Ion Mobility-Mass Spectrometry
复制标题

离子淌度质谱中代谢物碰撞截面值的大规模预测。

DOI:
10.1021/acs.analchem.6b03091
复制
发表时间:
2016-11-15
影响因子:
7.4
通讯作者:
Zhu, Zheng-Jiang
Zhu, Zheng-Jiang
中科院分区:
化学1区
文献类型:
--
作者:
Zhou, Zhiwei;Shen, Xiaotao;Zhu, Zheng-Jiang

文献摘要

被引文献

相似文献

代谢组学的快速发展极大地促进了健康和疾病相关研究。然而,代谢物鉴定仍然是非靶向代谢组学的主要分析挑战。虽然使用离子迁移率-质谱(IM-MS)中获得的碰撞截面(CCS)值有效地提高了代谢物的鉴别置信度,但其受到代谢物可用CCS值数量有限的限制。在这里,我们展示了使用一种称为支持向量回归(SVR)的机器学习算法来开发一种预测方法,该方法利用14种常见的分子描述符来预测代谢物的CCS值。在这项工作中,我们首先通过实验测量了氮缓冲气体中类似于400种代谢物的CCS值(Omega(N2)),并将这些值作为训练数据来优化预测方法。使用一组独立的代谢物对该方法的高预测精度进行了外部验证,其中位相对误差(MRE)接近3%,优于传统的理论计算。使用基于SVR的预测方法,在人类代谢物组数据库(HMDB)中生成35 203种代谢物的大规模预测CCS数据库。对于每种代谢产物,预测了5种不同的阳离子加合物和阴离子加合物,共占176 015 CCS值。最后,使用真实的生物样品证明了改进的代谢物鉴定准确度。总之,我们的研究结果证明,基于支持向量回归机的预测方法可以准确地预测氮CCS值(欧米茄(N2))的代谢产物的分子描述符,并有效地提高识别的准确性和效率在非靶向代谢组学。预测的CCS数据库,即MetCCS,可在互联网上免费获得。
The rapid development of metabolomics has significantly advanced health and disease related research. However, metabolite identification remains a major analytical challenge for untargeted metabolomics. While the use of collision cross-section (CCS) values obtained in ion mobility-mass spectrometry (IM-MS) effectively increases identification confidence of metabolites, it is restricted by the limited number of available CCS values for metabolites. Here, we demonstrated the use of a machine-learning algorithm called support vector regression (SVR) to develop a prediction method that utilized 14 common molecular descriptors to predict CCS values for metabolites. In this work, we first experimentally measured CCS values (Omega(N2)) of similar to 400 metabolites in nitrogen buffer gas and used these values as training data to optimize the prediction method. The high prediction precision of this method was externally validated using an independent set of metabolites with a median relative error (MRE) of similar to 3%, better than conventional theoretical calculation. Using the SVR based prediction method, a large-scale predicted CCS database was generated for 35 203 metabolites in the Human Metabolome Database (HMDB). For each metabolite, five different ion adducts in positive and negative modes were predicted, accounting for 176 015 CCS values in total. Finally, improved metabolite identification accuracy was demonstrated using real biological samples. Conclusively, our results proved that the SVR based prediction method can accurately predict nitrogen CCS values (Omega(N2)) of metabolites from molecular descriptors and effectively improve identification accuracy and efficiency in untargeted metabolomics. The predicted CCS database, namely, MetCCS, is freely available on the Internet.