A k-mer-based method for the identification of phenotype-associated genomic biomarkers and predicting phenotypes of sequenced bacteria

A k-mer-based method for the identification of phenotype-associated genomic biomarkers and predicting phenotypes of sequenced bacteria
复制标题

DOI:
10.1371/journal.pcbi.1006434
复制
发表时间:
2018-10-01
影响因子:
4.3
通讯作者:
Remm, Maido
Remm, Maido
中科院分区:
生物学2区
文献类型:
--
作者:
Aun, Erki;Brauer, Age;Remm, Maido

文献摘要

被引文献

相似文献

我们已经开发了一种易于使用和记忆高效的方法称为PhenotypeSeeker,其(a)鉴定表型特异性k-mer,(B)生成用于预测给定表型的基于k-mer的统计模型,以及(c)从给定细菌分离株的测序数据预测表型。该方法在167株肺炎克雷伯菌分离株(毒力)、200株铜绿假单胞菌分离株(环丙沙星耐药)和459株艰难梭菌分离株(阿奇霉素耐药)上进行了验证。从这些数据集训练的表型预测模型在K. pneumoniae测试集上为0.88,铜绿假单胞菌测试集上为0.97。difficile测试集。组装序列和原始测序数据的F1度量是相同的;然而,从组装的基因组构建模型要快得多。在这些数据集上,如果使用组装的基因组,则在中等范围的Linux服务器上构建模型对于每个表型需要大约3至5小时,如果使用原始测序数据,则对于每个表型需要10小时。从组装的基因组进行表型预测,每个分离株所需时间不到1秒。因此,PhenotypeSeeker应该非常适合从大型测序数据集预测表型。PhenotypeSeeker是用Python编程语言实现的,是开源软件,可在GitHub(https://github.com/bioinfo-ut/PhenotypeSeeker/)上获得。
We have developed an easy-to-use and memory-efficient method called PhenotypeSeeker that (a) identifies phenotype-specific k-mers, (b) generates a k-mer-based statistical model for predicting a given phenotype and (c) predicts the phenotype from the sequencing data of a given bacterial isolate. The method was validated on 167 Klebsiella pneumoniae isolates (virulence), 200 Pseudomonas aeruginosa isolates (ciprofloxacin resistance) and 459 Clostridium difficile isolates (azithromycin resistance). The phenotype prediction models trained from these datasets obtained the Fl-measure of 0.88 on the K. pneumoniae test set, 0.88 on the P. aeruginosa test set and 0.97 on the C. difficile test set. The F1-measures were the same for assembled sequences and raw sequencing data; however, building the model from assembled genomes is significantly faster. On these datasets, the model building on a mid-range Linux server takes approximately 3 to 5 hours per phenotype if assembled genomes are used and 10 hours per phenotype if raw sequencing data are used. The phenotype prediction from assembled genomes takes less than one second per isolate. Thus, PhenotypeSeeker should be well-suited for predicting phenotypes from large sequencing datasets. PhenotypeSeeker is implemented in Python programming language, is opensource software and is available at GitHub (https://github.com/bioinfo-ut/PhenotypeSeeker/).