A genetic algorithm-Bayesian network approach for the analysis of metabolomics and spectroscopic data: application to the rapid identification of Bacillus spores and classification of Bacillus species.

A genetic algorithm-Bayesian network approach for the analysis of metabolomics and spectroscopic data: application to the rapid identification of Bacillus spores and classification of Bacillus species.
复制标题

DOI:
10.1186/1471-2105-12-33
复制
发表时间:
2011-01-26
期刊:
影响因子:
3
通讯作者:
Goodacre R
Goodacre R
中科院分区:
生物学4区
文献类型:
--
作者:
Correa E;Goodacre R

文献摘要

参考文献

被引文献

相似文献

芽孢杆菌孢子的快速鉴定和细菌鉴定是至关重要的,因为它们在食物中毒、发病机制和作为潜在生物战剂的用途方面具有重要意义。许多自动化分析技术,如居里点热解质谱(Py-MS)已被用于鉴定细菌孢子,从而使用大量的分析数据。我们分析了36种不同的有氧内孢子形成细菌菌株的Py-MS数据,包括7种不同的物种。这些细菌在营养琼脂上无菌培养,用居里点Py-MS分析了营养生物量和孢子。我们开发了一种新的遗传算法-贝叶斯网络算法,该算法可以准确识别并选择一小部分关键相关质谱(生物标志物)进行进一步分析。一旦确定,这一相关生物标志物子集随后被用于成功识别芽孢杆菌孢子,并通过专门为这一简化特征集构建的贝叶斯网络模型识别芽孢杆菌物种。这个最终的紧凑贝叶斯网络分类模型是简约的,计算速度快,它的图形可视化可以很容易地解释选定的生物标志物之间的概率关系。此外,我们将遗传算法-贝叶斯网络方法选择的特征与偏最小二乘-判别分析(PLS-DA)选择的特征进行了比较。分类精度结果表明,GA-BN选择的特征集远优于PLS-DA。
The rapid identification of Bacillus spores and bacterial identification are paramount because of their implications in food poisoning, pathogenesis and their use as potential biowarfare agents. Many automated analytical techniques such as Curie-point pyrolysis mass spectrometry (Py-MS) have been used to identify bacterial spores giving use to large amounts of analytical data. This high number of features makes interpretation of the data extremely difficult We analysed Py-MS data from 36 different strains of aerobic endospore-forming bacteria encompassing seven different species. These bacteria were grown axenically on nutrient agar and vegetative biomass and spores were analyzed by Curie-point Py-MS. We develop a novel genetic algorithm-Bayesian network algorithm that accurately identifies sand selects a small subset of key relevant mass spectra (biomarkers) to be further analysed. Once identified, this subset of relevant biomarkers was then used to identify Bacillus spores successfully and to identify Bacillus species via a Bayesian network model specifically built for this reduced set of features. This final compact Bayesian network classification model is parsimonious, computationally fast to run and its graphical visualization allows easy interpretation of the probabilistic relationships among selected biomarkers. In addition, we compare the features selected by the genetic algorithm-Bayesian network approach with the features selected by partial least squares-discriminant analysis (PLS-DA). The classification accuracy results show that the set of features selected by the GA-BN is far superior to PLS-DA.
DOI: 10.1002/cem.785
发表时间: 2003-03-01
影响因子: 2.4
作者:
Barker, M;Rayens, W
通讯作者: Rayens, W
DOI: 10.1128/jb.00282-07
发表时间: 2007-07-01
影响因子: 3.2
作者:
Huang, Shu-Shi;Chen, De;Li, Yong-Qing
通讯作者: Li, Yong-Qing
DOI: 10.1366/0003702924125609
发表时间: 1992-02-01
影响因子: 3.5
作者:
GHIAMATI, E;MANOHARAN, R;SPERRY, JF
通讯作者: SPERRY, JF
DOI: 10.1021/ac990661i
发表时间: 2000-01-01
影响因子: 7.4
作者:
Goodacre, R;Shann, B;Logan, NA
通讯作者: Logan, NA
DOI: 10.1023/a:1001800507443
发表时间: 1999-05-01
影响因子: 2.6
作者:
Atrih, A;Foster, SJ
通讯作者: Foster, SJ