Improved Discrimination of Disease States Using Proteomics Data with the Updated Aristotle Classifier.

Improved Discrimination of Disease States Using Proteomics Data with the Updated Aristotle Classifier.
复制标题

使用更新的亚里士多德分类器,使用蛋白质组学数据改善了疾病状态的歧视。

DOI:
10.1021/acs.jproteome.1c00066
复制
发表时间:
2021-05-07
影响因子:
4.4
通讯作者:
Desaire H
Desaire H
中科院分区:
生物学2区
文献类型:
--
作者:
Hua D;Desaire H

文献摘要

参考文献

被引文献

相似文献

来自组学研究的质谱学数据集是区分疾病患者和识别生物标记物的最佳信息源。在每一次分析中可以查询到数千种蛋白质或内源性代谢物,其丰度跨越几个数量级。有效利用这些数据准确识别疾病状态的机器学习工具需求很高。虽然质谱学数据集蕴含着丰富的潜在有用信息,但有效利用数据可能具有挑战性,因为数据集中缺少条目,而且样本数量通常比特征数量少得多,这两个挑战使机器学习变得困难。为了解决这个问题,我们修改了一个新的监督分类工具--亚里士多德分类器,这样就可以更好地利用组学数据集来识别疾病状态。优化的分类器AC.2021是在多个数据集上与其前身和两个领先的监督分类工具支持向量机(SVM)和XGBoost进行基准测试的。新的分类器AC.2021在使用蛋白质组学数据进行的多项测试中表现优于现有工具。这里提供的分类器的基本代码对于希望在使用他们的组学数据集来识别疾病状态时提高分类精度的研究人员是有用的。
Mass spectrometry data sets from ‘omics studies are an optimal information source for discriminating patients with disease and identifying biomarkers. Thousands of proteins or endogenous metabolites can be queried in each analysis, spanning several orders of magnitude in abundance. Machine learning tools that effectively leverage these data to accurately identify disease states are in high demand. While mass spectrometry data sets are rich with potentially useful information, using the data effectively can be challenging because of missing entries in the data sets and because the number of samples is typically much smaller than the number of features, two challenges that make machine learning difficult. To address this problem, we have modified a new supervised classification tool, the Aristotle Classifier, so that ‘omics data sets can be better leveraged for identifying disease states. The optimized classifier, AC.2021, is benchmarked on multiple data sets against its predecessor and two leading supervised classification tools, Support Vector Machine (SVM) and XGBoost. The new classifier, AC.2021, outperformed existing tools on multiple tests using proteomics data. The underlying code for the classifier, provided herein, would be useful for researchers who desire improved classification accuracy when using their ‘omics data sets to identify disease states.
DOI: 10.1021/acs.jproteome.0c00663
发表时间: 2021-01-01
影响因子: 4.4
作者:
Manzi, Malena;Palazzo, Martin;Monge, Maria Eugenia
通讯作者: Monge, Maria Eugenia
DOI: 10.1021/acs.analchem.9b01606
发表时间: 2019-09-03
影响因子: 7.4
作者:
Hua, David;Patabandige, Milani Wijeweera;Desaire, Heather
通讯作者: Desaire, Heather
DOI: 10.1021/acs.jproteome.0c00247
发表时间: 2020-08-07
影响因子: 4.4
作者:
Li, Jiankang;Duan, Wenting;Lu, Tingli
通讯作者: Lu, Tingli
shot弹枪蛋白质组学中的基于等速标记的相对定量。
DOI: 10.1021/pr500880b
发表时间: 2014-12-05
影响因子: 4.4
作者:
Rauniyar, Navin;Yates, John R., III
通讯作者: Yates, John R., III
DOI: 10.3233/jad-201318
发表时间: 2021
影响因子: 4
作者:
Khan, Mostafa J.;Desaire, Heather;Lopez, Oscar L.;Kamboh, M. Ilyas;Robinson, Rena A. S.
通讯作者: Robinson, Rena A. S.