MetaBinG2: a fast and accurate metagenomic sequence classification system for samples with many unknown organisms.

MetaBinG2: a fast and accurate metagenomic sequence classification system for samples with many unknown organisms.
复制标题

MetaBinG2:快速准确的宏基因组序列分类系统,适用于含有许多未知生物的样品

DOI:
10.1186/s13062-018-0220-y
复制
发表时间:
2018-08-22
期刊:
影响因子:
5.5
通讯作者:
Wei C
Wei C
中科院分区:
生物学2区
文献类型:
--
作者:
Qiao Y;Jia B;Hu Z;Sun C;Xiang Y;Wei C

文献摘要

参考文献

相似文献

背景宏基因组序列分类方法很多,但大多数方法都依赖于已知生物的基因组序列。一个大部分的测序序列可能被归类为未知的,这大大削弱了我们对整个sample.ResultHere的理解,我们提出MetaBinG 2,宏基因组序列分类的快速方法,特别是对于大量的未知生物体的样品。MetaBinG 2基于序列合成,并使用GPU来加速其速度。一百万个100 bp的Illumina序列可以在大约1分钟内在计算机上用一个GPU卡进行分类。我们通过将MetaBinG 2与多种流行的现有方法进行比较来评估MetaBinG 2。然后将MetaBinG 2应用于CAMDA数据分析竞赛提供的MetaBinG Inter-City Challenge数据集,比较了不同城市不同公共场所环境样品的群落组成结构。结论与现有方法相比,MetaBinG 2快速准确,特别是对于那些含有大量未知生物的样品。ReviewersThis article was reviewed by Drs. Eran Elhaik,Nicolas Rascovan,和谢尔盖·曼古尔
BackgroundMany methods have been developed for metagenomic sequence classification, and most of them depend heavily on genome sequences of the known organisms. A large portion of sequencing sequences may be classified as unknown, which greatly impairs our understanding of the whole sample.ResultHere we present MetaBinG2, a fast method for metagenomic sequence classification, especially for samples with a large number of unknown organisms. MetaBinG2 is based on sequence composition, and uses GPUs to accelerate its speed. A million 100 bp Illumina sequences can be classified in about 1 min on a computer with one GPU card. We evaluated MetaBinG2 by comparing it to multiple popular existing methods. We then applied MetaBinG2 to the dataset of MetaSUB Inter-City Challenge provided by CAMDA data analysis contest and compared community composition structures for environmental samples from different public places across cities.ConclusionCompared to existing methods, MetaBinG2 is fast and accurate, especially for those samples with significant proportions of unknown organisms.ReviewersThis article was reviewed by Drs. Eran Elhaik, Nicolas Rascovan, and Serghei Mangul.
DOI: 10.1186/s12859-015-0788-5
发表时间: 2015-11-04
期刊: BMC bioinformatics
影响因子: 3
作者:
Peabody MA;Van Rossum T;Lo R;Brinkman FS
通讯作者: Brinkman FS
MetaBinG:使用 GPU 加速宏基因组序列分类
DOI: 10.1371/journal.pone.0025353
发表时间: 2011
期刊: PloS one
影响因子: 3.7
作者:
Jia P;Xuan L;Liu L;Wei C
通讯作者: Wei C
DOI: 10.1007/s00018-015-2004-1
发表时间: 2015-11
期刊: Cellular and molecular life sciences : CMLS
影响因子: --
作者:
Garza DR;Dutilh BE
通讯作者: Dutilh BE
克拉克:使用判别性k-mers对宏基因组和基因组序列进行快速准确分类。
DOI: 10.1186/s12864-015-1419-2
发表时间: 2015-03-25
期刊: BMC genomics
影响因子: 4.4
作者:
Ounit R;Wanamaker S;Close TJ;Lonardi S
通讯作者: Lonardi S
DOI: 10.1093/nar/gks828
发表时间: 2013-01-07
影响因子: 14.9
作者:
Liu J;Wang H;Yang H;Zhang Y;Wang J;Zhao F;Qi J
通讯作者: Qi J