Incorporating microbial community data with machine learning techniques to predict feed substrates in microbial fuel cells

Incorporating microbial community data with machine learning techniques to predict feed substrates in microbial fuel cells
复制标题

将微生物群落数据与机器学习技术相结合来预测微生物燃料电池中的进料底物

DOI:
10.1016/j.bios.2019.03.021
复制
发表时间:
2019-05-15
影响因子:
12.6
通讯作者:
Liu, Hong
Liu, Hong
中科院分区:
工程技术1区
文献类型:
--
作者:
Cai, Wenfang;Lesnik, Keaton Larson;Liu, Hong

文献摘要

被引文献

相似文献

混合物种生物技术中发生的复杂相互作用,包括生物传感器,阻碍了化学检测的特异性。这种特异性的缺乏限制了生物传感器的应用,例如必须确定未知饲料底物的应用。基因组数据和成熟的数据挖掘技术的应用可以克服这些限制,促进工程发展。在本研究中,对从不同实验室环境中收集的三种不同底物类型(醋酸盐、碳水化合物和废水)的69个样品进行了评估,以确定从产生的微生物群落中识别饲料底物的能力。对具有4个不同输入变量的6种机器学习算法进行了训练,并对它们从基因组数据集预测饲料底物的能力进行了评估。在门和科分类水平的数据集上训练的NNET分别获得了93+/-6%和92+/-5%的最高准确率。卡伯值分别为0.87+/-0.10、0.86+/-0.09。在使用的六种算法中,有四种保持了80%以上的准确率和高于0.66的kappa值。不同的测序方法(Roche 454或Illumina测序)不影响所有算法的准确性,但在门水平上的支持向量机除外。在NMDS压缩的数据集上训练的所有算法的准确率都在80%以上,而在PCoA压缩的数据集上训练的模型的准确率降低了10%-30%。这些结果表明,将微生物群落数据与机器学习算法相结合可以用于饲料底物的预测和基于MFC的生物传感器信号特异性的潜在改进,为机器学习技术在生物技术领域具有实质性的实际应用提供了一种新的用途。
The complicated interactions that occur in mixed-species biotechnologies, including biosensors, hinder chemical detection specificity. This lack of specificity limits applications in which biosensors may be deployed, such as those where an unknown feed substrate must be determined. The application of genomic data and well-developed data mining technologies can overcome these limitations and advance engineering development. In the present study, 69 samples with three different substrate types (acetate, carbohydrates and wastewater) collected from various laboratory environments were evaluated to determine the ability to identify feed substrates from the resultant microbial communities. Six machine learning algorithms with four different input variables were trained and evaluated on their ability to predict feed substrate from genomic datasets. The highest accuracies of 93 +/- 6% and 92 +/- 5% were obtained using NNET trained on datasets classified at the phylum and family taxonomic level, respectively. These accuracies corresponded to kappa values of 0.87 +/- 0.10, 0.86 +/- 0.09, respectively. Four out of six of the algorithms used maintained accuracies above 80% and kappa values higher than 0.66. Different sequencing method (Roche 454 or Illumina sequencing) did not affect the accuracies of all algorithms, except SVM at the phylum level. All algorithms trained on NMDS-compressed datasets obtained accuracies over 80%, while models trained on PCoA-compressed datasets presented a 10-30% reduction in accuracy. These results suggest that incorporating microbial community data with machine learning algorithms can be used for the prediction of feed substrate and for the potential improvement of MFC-based biosensor signal specificity, providing a new use of machine learning techniques that has substantial practical applications in biotechnological fields.