Benchmark of structured machine learning methods for microbial identification from mass-spectrometry data

Benchmark of structured machine learning methods for microbial identification from mass-spectrometry data
复制标题

从质谱数据识别微生物的结构化机器学习方法的基准

DOI:
--
复制
发表时间:
2015
期刊:
arXiv.org
影响因子:
--
通讯作者:
Jean
Jean
中科院分区:
--
文献类型:
--
作者:
K. Vervier;P. Mahé;Jean;Jean

文献摘要

被引文献

相似文献

微生物鉴定是微生物学的核心问题,特别是在传染病诊断和工业质量控制领域。物种的概念与生物学和临床分类的概念密切相关,其中物种之间的接近程度通常以进化距离和/或临床表型来衡量。令人惊讶的是,这种众所周知的分层结构提供的信息很少被基于机器学习的自动微生物识别系统使用。最近提出了结构化机器学习方法,用于考虑嵌入在层次结构中的结构,并将其用作额外的先验信息,因此可以改进微生物鉴定系统。我们在一个新的基质辅助激光解吸/电离飞行时间质谱(MALDI-TOF MS)数据集上测试和比较了几种最先进的机器学习方法用于微生物鉴定。我们在基准标准和结构化方法中包括了在学习过程中利用底层层次结构知识的方法。我们的研究结果表明,虽然有些方法比其他方法表现更好,结构化的方法并不总是比他们的“平面”同行表现更好。我们假设,这部分是由于这样一个事实,即标准的方法已经达到了很高的准确性,在这种情况下,他们主要混淆物种接近对方的树,使用已知的层次结构是没有帮助的情况下。
Microbial identification is a central issue in microbiology, in particular in the fields of infectious diseases diagnosis and industrial quality control. The concept of species is tightly linked to the concept of biological and clinical classification where the proximity between species is generally measured in terms of evolutionary distances and/or clinical phenotypes. Surprisingly, the information provided by this well-known hierarchical structure is rarely used by machine learning-based automatic microbial identification systems. Structured machine learning methods were recently proposed for taking into account the structure embedded in a hierarchy and using it as additional a priori information, and could therefore allow to improve microbial identification systems. We test and compare several state-of-the-art machine learning methods for microbial identification on a new Matrix-Assisted Laser Desorption/Ionization Time-of-Flight mass spectrometry (MALDI-TOF MS) dataset. We include in the benchmark standard and structured methods, that leverage the knowledge of the underlying hierarchical structure in the learning process. Our results show that although some methods perform better than others, structured methods do not consistently perform better than their "flat" counterparts. We postulate that this is partly due to the fact that standard methods already reach a high level of accuracy in this context, and that they mainly confuse species close to each other in the tree, a case where using the known hierarchy is not helpful.