Chemical Class Prediction of Unknown Biomolecules Using Ion Mobility-Mass Spectrometry and Machine Learning: Supervised Inference of Feature Taxonomy from Ensemble Randomization.

Chemical Class Prediction of Unknown Biomolecules Using Ion Mobility-Mass Spectrometry and Machine Learning: Supervised Inference of Feature Taxonomy from Ensemble Randomization.
复制标题

DOI:
10.1021/acs.analchem.0c02137
复制
发表时间:
2020-08-04
影响因子:
7.4
通讯作者:
McLean, John A.
McLean, John A.
中科院分区:
化学1区
文献类型:
--
作者:
Picache, Jaqueline A.;May, Jody C.;McLean, John A.

文献摘要

参考文献

被引文献

相似文献

这项工作提出了一种机器学习算法,被称为监督推理的特征分类从包围随机化(SIFTER),它支持识别来自非目标离子迁移率质谱(IM-MS)实验的功能。SIFTER利用随机森林机器学习对来自IM-MS的三个分析测量值(碰撞截面(z/Ω),质荷比(m/z)和质量亏损(Δm))进行分类,将未知特征分类为化学王国,超类,类和子类。这些分类中的每一个都被分配了计算的概率以及具有相关联的概率的替代分类。优化后,针对训练集中未使用的一组分子测试SIFTER。发现正确分类所有四个分类类别的平均成功率> 99%。从复杂的生物基质中检测到的分子特征的分析,而不是在训练集中使用,产生了较低的成功率,其中所有四个类别都正确预测了约80%的化合物。这种性能下降部分是由于所有潜在分类类别的训练集不完整,但也是由于随机森林算法中的最近邻偏差。正在进行的努力集中在通过扩展用于训练的经验数据集以及改进核心算法来提高SIFTER的类预测准确性。
This work presents a machine learning algorithm referred to as the Supervised Inference of Feature Taxonomy from Ensemble Randomization (SIFTER), which supports the identification of features derived from untargeted ion mobility-mass spectrometry (IM-MS) experiments. SIFTER utilizes random forest machine learning on three analytical measurements derived from IM-MS (collision cross section (z/Ω), mass-to-charge (m/z), and mass defect (Δm)) to classify unknown features into a taxonomy of chemical kingdom, super class, class, and subclass. Each of these classifications is assigned a calculated probability as well as alternate classifications with associated probabilities. After optimization, SIFTER was tested against a set of molecules not used in the training set. The average success rate in classifying all four taxonomy categories correctly was found to be >99%. Analysis of molecular features detected from a complex biological matrix and not used in the training set yielded a lower success rate where all four categories were correctly predicted for ~80% of the compounds. This decline in performance is in part due to incompleteness of the training set across all potential taxonomic categories, but also resulting from a nearest neighbor bias in the random forest algorithm. Ongoing efforts are focused on improving the class prediction accuracy of SIFTER through expansion of empirical datasets used for training as well as improvements to the core algorithm.
DOI: 10.1186/s13321-018-0324-5
发表时间: 2019-01-05
影响因子: 8.6
作者:
Djoumbou-Feunang, Yannick;Fiamoncini, Jarlei;Wishart, David S.
通讯作者: Wishart, David S.
使用秀丽隐杆线虫自动表型分析对化学品毒性进行分类和预测
DOI: 10.1186/s40360-018-0208-3
发表时间: 2018-04-18
影响因子: 2.9
作者:
Gao S;Chen W;Zeng Y;Jing H;Zhang N;Flavel M;Jois M;Han JJ;Xian B;Li G
通讯作者: Li G
DOI: 10.1146/annurev-anchem-071015-041734
发表时间: 2016-06-12
期刊: Annual review of analytical chemistry (Palo Alto, Calif.)
影响因子: --
作者:
May JC;McLean JA
通讯作者: McLean JA
DOI: 10.1093/bioinformatics/btz954
发表时间: 2020-04-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Baranwal, Mayank;Magner, Abram;Hero, Alfred O.
通讯作者: Hero, Alfred O.
DOI: 10.1021/acs.analchem.6b03091
发表时间: 2016-11-15
影响因子: 7.4
作者:
Zhou, Zhiwei;Shen, Xiaotao;Zhu, Zheng-Jiang
通讯作者: Zhu, Zheng-Jiang