Analysis and comparison of very large metagenomes with fast clustering and functional annotation.

Analysis and comparison of very large metagenomes with fast clustering and functional annotation.
复制标题

DOI:
10.1186/1471-2105-10-359
复制
发表时间:
2009-10-28
期刊:
影响因子:
3
通讯作者:
Li W
Li W
中科院分区:
生物学4区
文献类型:
--
作者:
Li W

文献摘要

参考文献

被引文献

相似文献

宏基因组学的显著进步对数据分析提出了重大的新挑战。宏基因组数据集(宏基因组)是来自特定环境中的匿名物种的测序读数的大集合。非常大的宏基因组的计算分析是非常耗时的,并且在这些宏基因组中通常有许多未被充分利用的新序列。可用宏基因组的数量正在迅速增加,因此快速有效的宏基因组比较方法需求量很大。新的宏基因组数据分析方法快速分析多个宏基因组与聚类和注释管道(RAMMCAP)开发使用超快速序列聚类算法,快速蛋白质家族注释工具,和一种新的统计宏基因组比较方法,采用独特的图形界面。RAMMCAP处理非常大的数据集,只有适度的计算工作。它识别可能包括新基因家族的原始读段簇和蛋白簇,并使用RAMMCAP计算的簇或功能注释比较宏基因组。在这项研究中,RAMMCAP被应用于两个最大的宏基因组收集,“全球海洋采样”和“九个生物群的宏基因组分析”。RAMMCAP是一种非常快速的方法,可以在数百个CPU小时内聚类和注释一百万个宏基因组读数。可从以下网站获得。
The remarkable advance of metagenomics presents significant new challenges in data analysis. Metagenomic datasets (metagenomes) are large collections of sequencing reads from anonymous species within particular environments. Computational analyses for very large metagenomes are extremely time-consuming, and there are often many novel sequences in these metagenomes that are not fully utilized. The number of available metagenomes is rapidly increasing, so fast and efficient metagenome comparison methods are in great demand. The new metagenomic data analysis method Rapid Analysis of Multiple Metagenomes with a Clustering and Annotation Pipeline (RAMMCAP) was developed using an ultra-fast sequence clustering algorithm, fast protein family annotation tools, and a novel statistical metagenome comparison method that employs a unique graphic interface. RAMMCAP processes extremely large datasets with only moderate computational effort. It identifies raw read clusters and protein clusters that may include novel gene families, and compares metagenomes using clusters or functional annotations calculated by RAMMCAP. In this study, RAMMCAP was applied to the two largest available metagenomic collections, the "Global Ocean Sampling" and the "Metagenomic Profiling of Nine Biomes". RAMMCAP is a very fast method that can cluster and annotate one million metagenomic reads in only hundreds of CPU hours. It is available from .
Metasim:用于基因组学和元基因组学的测序模拟器。
DOI: 10.1371/journal.pone.0003373
发表时间: 2008-10-08
期刊: PLOS ONE
影响因子: 3.7
作者:
Richter, Daniel C.;Ott, Felix;Auch, Alexander F.;Schmid, Ramona;Huson, Daniel H.
通讯作者: Huson, Daniel H.
DOI: 10.1093/dnares/dsn027
发表时间: 2008-12
期刊: DNA RESEARCH
影响因子: 4.1
作者:
Noguchi, Hideki;Taniguchi, Takeaki;Itoh, Takehiko
通讯作者: Itoh, Takehiko
DOI: 10.1186/1471-2105-7-162
发表时间: 2006-03-20
期刊: BMC bioinformatics
影响因子: 3
作者:
Rodriguez-Brito B;Rohwer F;Edwards RA
通讯作者: Edwards RA
DOI: 10.1038/nmeth1043
发表时间: 2007-06-01
期刊: NATURE METHODS
影响因子: 48
作者:
Mavromatis, Konstantinos;Ivanova, Natalia;Kyrpides, Nikos C.
通讯作者: Kyrpides, Nikos C.
DOI: 10.1093/bioinformatics/17.3.282
发表时间: 2001-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, WZ;Jaroszewski, L;Godzik, A
通讯作者: Godzik, A