Improved metagenome assemblies and taxonomic binning using long-read circular consensus sequence data.

Improved metagenome assemblies and taxonomic binning using long-read circular consensus sequence data.
复制标题

DOI:
10.1038/srep25373
复制
发表时间:
2016-05-09
期刊:
影响因子:
4.6
通讯作者:
Pope PB
Pope PB
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Frank JA;Pan Y;Tooming-Klunderud A;Eijsink VGH;McHardy AC;Nederbragt AJ;Pope PB

文献摘要

被引文献

相似文献

DNA组装是用于研究微生物群落结构和功能的宏基因组管道中的核心方法步骤。在这里,我们研究了太平洋生物科学公司的长和高准确度的环状共识测序(CCS)读段的宏基因组项目的效用。我们比较了PacBio CCS和Illumina HiSeq数据的应用和性能,以及使用代表复杂微生物群落的宏基因组样本的组装和分类分箱算法。八个SMRT细胞从沼气反应器微生物组样本中产生了约94 Mb的CCS读数,平均长度为1319 nt,准确度为99.7%。CCS数据组装产生了与从相同样品产生的约190倍更大的HiSeq数据集(约18 Gb)组装的那些(即约62%的总重叠群)相比数量的大于1 kb的大重叠群。使用PacBio CCS和HiSeq重叠群的混合组装产生了组装统计的改进,包括平均重叠群长度和大重叠群数量的增加。CCS数据的纳入产生了显着的增强,在分类分箱和基因组重建的两个占主导地位的基因组类型,组装和分箱不好单独使用HiSeq数据。总的来说,这些结果说明了PacBio CCS读数在某些宏基因组学应用中的价值。
DNA assembly is a core methodological step in metagenomic pipelines used to study the structure and function within microbial communities. Here we investigate the utility of Pacific Biosciences long and high accuracy circular consensus sequencing (CCS) reads for metagenomic projects. We compared the application and performance of both PacBio CCS and Illumina HiSeq data with assembly and taxonomic binning algorithms using metagenomic samples representing a complex microbial community. Eight SMRT cells produced approximately 94 Mb of CCS reads from a biogas reactor microbiome sample that averaged 1319 nt in length and 99.7% accuracy. CCS data assembly generated a comparative number of large contigs greater than 1 kb, to those assembled from a ~190x larger HiSeq dataset (~18 Gb) produced from the same sample (i.e approximately 62% of total contigs). Hybrid assemblies using PacBio CCS and HiSeq contigs produced improvements in assembly statistics, including an increase in the average contig length and number of large contigs. The incorporation of CCS data produced significant enhancements in taxonomic binning and genome reconstruction of two dominant phylotypes, which assembled and binned poorly using HiSeq data alone. Collectively these results illustrate the value of PacBio CCS reads in certain metagenomics applications.