Keeping up with the genomes: efficient learning of our increasing knowledge of the tree of life.

Keeping up with the genomes: efficient learning of our increasing knowledge of the tree of life.
复制标题

DOI:
10.1186/s12859-020-03744-7
复制
发表时间:
2020-09-21
期刊:
影响因子:
3
通讯作者:
Rosen G
Rosen G
中科院分区:
生物学4区
文献类型:
--
作者:
Zhao Z;Cristian A;Rosen G

文献摘要

参考文献

被引文献

相似文献

对于当前的宏基因组分类器来说,要跟上基因组测序项目(如NCBI RefSeq细菌基因组数据库)产生的训练数据的步伐是一个计算上的挑战。当向训练数据中添加新的引用序列时,必须在所有数据上重新运行静态训练的分类器,从而导致效率极低。“增量学习”的丰富文献解决了更新现有分类器以适应新数据的需求,而与使用所有数据重新训练分类器相比,又不会牺牲太多的准确性。我们通过在渐进RefSeq快照上增量训练分类器并在:(a)所有已知的当前基因组(作为基础真理集)和(b)真实的实验宏基因组肠道样本上进行测试来演示分类如何随着时间的推移而改进。我们证明,随着分类器模型对基因组知识的增长,分类精度也会提高。概念验证naïve贝叶斯实现,当每年更新时,现在运行在1/4的非增量时间,没有准确性损失。很明显,分类法由于掌握了最新的知识而得到改进。因此,使分类器在计算上易于处理以跟上数据洪流是至关重要的。增量学习分类器可以有效地更新,而不需要重新处理,也不需要访问现有数据库,因此节省了存储和计算资源。
It is a computational challenge for current metagenomic classifiers to keep up with the pace of training data generated from genome sequencing projects, such as the exponentially-growing NCBI RefSeq bacterial genome database. When new reference sequences are added to training data, statically trained classifiers must be rerun on all data, resulting in a highly inefficient process. The rich literature of “incremental learning” addresses the need to update an existing classifier to accommodate new data without sacrificing much accuracy compared to retraining the classifier with all data. We demonstrate how classification improves over time by incrementally training a classifier on progressive RefSeq snapshots and testing it on: (a) all known current genomes (as a ground truth set) and (b) a real experimental metagenomic gut sample. We demonstrate that as a classifier model’s knowledge of genomes grows, classification accuracy increases. The proof-of-concept naïve Bayes implementation, when updated yearly, now runs in 1/4th of the non-incremental time with no accuracy loss. It is evident that classification improves by having the most current knowledge at its disposal. Therefore, it is of utmost importance to make classifiers computationally tractable to keep up with the data deluge. The incremental learning classifier can be efficiently updated without the cost of reprocessing nor the access to the existing database and therefore save storage as well as computation resources.
DOI: 10.1186/s13059-017-1299-7
发表时间: 2017-09-21
期刊: Genome biology
影响因子: 12.3
作者:
McIntyre ABR;Ounit R;Afshinnekoo E;Prill RJ;Hénaff E;Alexander N;Minot SS;Danko D;Foox J;Ahsanuddin S;Tighe S;Hasan NA;Subramanian P;Moffat K;Levy S;Lonardi S;Greenfield N;Colwell RR;Rosen GL;Mason CE
通讯作者: Mason CE
DOI: 10.1093/bioinformatics/btt389
发表时间: 2013-09-15
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Ames SK;Hysom DA;Gardner SN;Lloyd GS;Gokhale MB;Allen JE
通讯作者: Allen JE
DOI: 10.1109/5326.983933
发表时间: 2001-11-01
期刊: IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS PART C-APPLICATIONS AND REVIEWS
影响因子: --
作者:
Polikar, R;Udpa, L;Honavar, V
通讯作者: Honavar, V
DOI: 10.1186/s12859-018-2182-6
发表时间: 2018-07-09
期刊: BMC bioinformatics
影响因子: 3
作者:
Fiannaca A;La Paglia L;La Rosa M;Lo Bosco G;Renda G;Rizzo R;Gaglio S;Urso A
通讯作者: Urso A
克拉克:使用判别性k-mers对宏基因组和基因组序列进行快速准确分类。
DOI: 10.1186/s12864-015-1419-2
发表时间: 2015-03-25
期刊: BMC genomics
影响因子: 4.4
作者:
Ounit R;Wanamaker S;Close TJ;Lonardi S
通讯作者: Lonardi S