The local-balanced model for improved machine learning outcomes on mass spectrometry data sets and other instrumental data.

The local-balanced model for improved machine learning outcomes on mass spectrometry data sets and other instrumental data.
复制标题

用于改进对质谱学数据集和其他仪器数据的机器学习结果的本地平衡模型。

DOI:
10.1007/s00216-020-03117-2
复制
发表时间:
2021-03
影响因子:
4.3
通讯作者:
Hua D
Hua D
中科院分区:
化学2区
文献类型:
--
作者:
Desaire H;Patabandige MW;Hua D

文献摘要

参考文献

被引文献

相似文献

当用质谱数据对生物样品进行分类时,一个统一的挑战是克服样品间变异性的障碍,使得可以识别组之间的差异,例如健康组和疾病组之间的差异。同样,当在相同条件下重新分析相同样品时,仪器信号可能波动超过10%。这种信号的不一致性给识别一组样品之间的细微差异带来了困难,并且它削弱了质谱分析师有效利用蛋白质组学、代谢组学、糖组学和成像等领域数据的能力。我们选择了糖组学、质谱成像和细菌分型领域中具有挑战性的数据集来研究组内信号变异性的问题,并采用了一种已有30年历史的统计方法来解决这个问题。解决方案“局部平衡模型”依赖于使用训练数据的平衡子集来对测试样本进行分类。基于IgG基糖肽的ESI-MS数据和内源性脂质的MALDI-MS成像数据以及细菌蛋白的MALDI-MS数据评估了该分析策略。两个初步的例子,非质谱数据集也包括显示MS分析领域外的方法的潜在的一般性。我们证明,这种方法是上级简单的归一化方法,可推广到多个质谱领域,并可能适用于物理和卫星成像等领域。在某些情况下,分类的改进可能是显著的,准确度从单独标准化的60%上升到本文所述的额外开发的90%以上。
One unifying challenge when classifying biological samples with mass spectrometry data is overcoming the obstacle of sample-to-sample variability so that differences between groups, such as between a healthy set and a disease set, can be identified. Similarly, when the same sample is re-analyzed under identical conditions, instrument signals can fluctuate by more than 10%. This signal inconsistency imposes difficulties in identifying subtle differences across a set of samples, and it weakens the mass spectrometrist’s ability to effectively leverage data in domains as diverse as proteomics, metabolomics, glycomics, and imaging. We selected challenging data sets in the fields of glycomics, mass spectrometry imaging, and bacterial typing to study the problem of within-group signal variability and adapted a 30 year old statistical approach to address the problem. The solution, “local-balanced model,” relies on using balanced subsets of training data to classify test samples. This analysis strategy was assessed on ESI-MS data of IgG-based glycopeptides and MALDI-MS imaging data of endogenous lipids, and MALDI-MS data of bacterial proteins. Two preliminary examples on non-mass spectrometry data sets are also included to show the potential generality of the method outside the field of MS analysis. We demonstrate that this approach is superior to simple normalization methods, generalizable to multiple mass spectrometry domains, and potentially appropriate in fields as diverse as physics and satellite imaging. In some cases, improvements in classification can be dramatic, with accuracy escalating from 60% with normalization alone to over 90% with the additional development described herein.
DOI: 10.1093/bioinformatics/btu022
发表时间: 2014-05-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Mahe, Pierre;Arsac, Maud;Veyrieras, Jean-Baptiste
通讯作者: Veyrieras, Jean-Baptiste
DOI: 10.1021/acs.analchem.8b05985
发表时间: 2019-05-07
影响因子: 7.4
作者:
Liu, Zhichao;Portero, Erika P.;Nemes, Peter
通讯作者: Nemes, Peter
DOI: 10.1021/acs.analchem.9b01606
发表时间: 2019-09-03
影响因子: 7.4
作者:
Hua, David;Patabandige, Milani Wijeweera;Desaire, Heather
通讯作者: Desaire, Heather
DOI: 10.1109/tsmcb.2008.2007853
发表时间: 2009-04-01
影响因子: --
作者:
Liu, Xu-Ying;Wu, Jianxin;Zhou, Zhi-Hua
通讯作者: Zhou, Zhi-Hua
DOI: 10.1109/tgrs.2008.916090
发表时间: 2008-06-01
影响因子: 8.2
作者:
Blanzieri, Enrico;Melgani, Farid
通讯作者: Melgani, Farid