A Machine-Learning Approach to Detecting Unknown Bacterial Serovars.

A Machine-Learning Approach to Detecting Unknown Bacterial Serovars.
复制标题

DOI:
10.1002/sam.10085
复制
发表时间:
2010-10
影响因子:
1.3
通讯作者:
Rajwa, Bartek
Rajwa, Bartek
中科院分区:
计算机科学4区
文献类型:
--
作者:
Akova, Ferit;Dundar, Murat;Davisson, V Jo;Hirleman, E Daniel;Bhunia, Arun K;Robinson, J Paul;Rajwa, Bartek

文献摘要

被引文献

相似文献

快速检测细菌病原体的技术对于确保食品供应至关重要。最近开发的用于实时识别多个菌落的光散射传感器在区分细菌培养物方面显示出很大的前景。该系统目前使用的分类方法依赖于监督学习。为了准确分类细菌病原体,训练库应该是详尽的,即,应该包括所有可能的病原体的样本。然而,现有的细菌血清型的绝对数量,更重要的是它们的高突变率的影响,将不允许实际和可管理的培训。在这项研究中,我们提出了一种贝叶斯方法来学习非穷举训练数据集,用于自动检测不匹配的细菌血清型,即,训练库中不存在样本的血清型。我们的工作的主要贡献是Wishart共轭先验定义类分布。这使我们能够利用从已知类中获得的先验信息来推断未知类。通过这种方式,我们识别新的信息值类别,并使用这些类别动态更新训练数据集,使其越来越具有样本群体的代表性。这导致分类器对未来样本具有改进的预测性能。我们在28类细菌数据集上评估了我们的方法,并在基准的26类字母识别数据集上进行了进一步验证。所提出的方法进行比较,对国家的最先进的涉及基于密度的方法和支持向量域描述,以及最近推出的贝叶斯方法的基础上模拟类。
Technologies for rapid detection of bacterial pathogens are crucial for securing the food supply. A light-scattering sensor recently developed for real-time identification of multiple colonies has shown great promise for distinguishing bacteria cultures. The classification approach currently used with this system relies on supervised learning. For accurate classification of bacterial pathogens, the training library should be exhaustive, i.e., should consist of samples of all possible pathogens. Yet, the sheer number of existing bacterial serovars and more importantly the effect of their high mutation rate would not allow for a practical and manageable training. In this study, we propose a Bayesian approach to learning with a nonexhaustive training dataset for automated detection of unmatched bacterial serovars, i.e., serovars for which no samples exist in the training library. The main contribution of our work is the Wishart conjugate priors defined over class distributions. This allows us to employ the prior information obtained from known classes to make inferences about unknown classes as well. By this means, we identify new classes of informational value and dynamically update the training dataset with these classes to make it increasingly more representative of the sample population. This results in a classifier with improved predictive performance for future samples. We evaluated our approach on a 28-class bacteria dataset and also on the benchmark 26-class letter recognition dataset for further validation. The proposed approach is compared against state-of-the-art involving density-based approaches and support vector domain description, as well as a recently introduced Bayesian approach based on simulated classes.