MinION™ nanopore sequencing of environmental metagenomes: a synthetic approach.

MinION™ nanopore sequencing of environmental metagenomes: a synthetic approach.
复制标题

DOI:
10.1093/gigascience/gix007
复制
发表时间:
2017-03-01
期刊:
影响因子:
9.2
通讯作者:
Franklin RB
Franklin RB
中科院分区:
生物学2区
文献类型:
--
作者:
Brown BL;Watson M;Minot SS;Rivera MC;Franklin RB

文献摘要

被引文献

相似文献

背景资料:环境宏基因组分析通常通过从全基因组测序或16 S扩增子序列分配分类和/或功能来完成。然而,这两种方法都受到阅读长度以及其他技术和生物因素的限制。基于纳米孔的测序平台MinION™产生长度≥1 × 104 bp的读段,可能提供更精确的分配,从而减轻从短读段确定宏基因组组成的一些固有限制。我们测试了MinION产生的序列数据的能力(R7.3流动池)在单一菌种运行和三种类型的低复杂性合成群落中正确分配分类:一种来自四个物种的等质量DNA混合物,一种相对罕见(1%),三种丰富(各33%)组分,以及来自交错代表的20种细菌菌株的基因组DNA的混合物。低复杂性社区的分类组成进行了评估,通过分析MinION序列数据与三种不同的生物信息学方法:Kraken,MG-RAST,和一个法典。结果如下:使用第5版试剂盒和化学方法从单菌株制备的文库中生成的长读段序列在原始MinION设备上运行,产生了少至224至多达3497个双向高质量(2D)读段,平均总研究长度为6000 bp。对于单菌株分析,通过不同方法将读数分配至正确属的范围为53.1%至99.5%,分配至正确种的范围为23.9%至99.5%,并且大多数错误分配的读数是密切相关的微生物。用相同设置测序的合成宏基因组产生了714个约5500 bp的高质量2D读段,其高达98%正确地分配到物种水平。使用版本6试剂盒和化学产生的合成宏基因组MinION文库产生899至3497个2D读数,平均长度为5700 bp,在物种水平上具有高达98%的分配准确度。观察到的“相等”和“稀有”合成文库的群落比例接近已知比例,在所有测试中偏离0.1%至10%。对于具有交错贡献的20个物种的模拟群落,测序运行检测到除了3个物种之外的所有物种(每个物种在总混合物中包括<0.05%的DNA),91%的读段被分配给正确的物种,93%的读段被分配给正确的属,并且>99%的读段被分配给正确的家族。结论:在目前的输出和序列质量水平下(合成宏基因组的2D读数略低于4 × 103),MinION测序后进行Kraken或One Codex分析有可能提供快速准确的宏基因组分析,其中该聚生体由有限数量的分类群组成。本研究中注意到的重要考虑因素包括:MinION平台对输入DNA质量的高灵敏度,文库和流动池之间测序结果的高变异性,以及每个分析限值的2D读数数量相对较少。总之,这些限制了对微生物聚生体的非常罕见的组分的检测,并且可能限制MinION用于测序高复杂性宏基因组群落的效用,其中预期有数千个分类群。此外,目前可用的数据分析工具的局限性表明,使用长读段表征微生物群落的分析方法有相当大的改进空间。然而,MinION生成的高质量读数的准确分类学分配接近99.5%,并且在大多数情况下,推断的社区结构反映了合成混合物的已知比例,这一事实证明,随着平台的不断发展和改进,进一步探索环境宏基因组学的实际应用。随着序列通量的进一步提高和错误率的降低,该平台显示出对更复杂微生物群落的组成和结构进行精确实时分析的巨大潜力。
Background: Environmental metagenomic analysis is typically accomplished by assigning taxonomy and/or function from whole genome sequencing or 16S amplicon sequences. Both of these approaches are limited, however, by read length, among other technical and biological factors. A nanopore-based sequencing platform, MinION™, produces reads that are ≥1 × 104 bp in length, potentially providing for more precise assignment, thereby alleviating some of the limitations inherent in determining metagenome composition from short reads. We tested the ability of sequence data produced by MinION (R7.3 flow cells) to correctly assign taxonomy in single bacterial species runs and in three types of low-complexity synthetic communities: a mixture of DNA using equal mass from four species, a community with one relatively rare (1%) and three abundant (33% each) components, and a mixture of genomic DNA from 20 bacterial strains of staggered representation. Taxonomic composition of the low-complexity communities was assessed by analyzing the MinION sequence data with three different bioinformatic approaches: Kraken, MG-RAST, and One Codex. Results: Long read sequences generated from libraries prepared from single strains using the version 5 kit and chemistry, run on the original MinION device, yielded as few as 224 to as many as 3497 bidirectional high-quality (2D) reads with an average overall study length of 6000 bp. For the single-strain analyses, assignment of reads to the correct genus by different methods ranged from 53.1% to 99.5%, assignment to the correct species ranged from 23.9% to 99.5%, and the majority of misassigned reads were to closely related organisms. A synthetic metagenome sequenced with the same setup yielded 714 high quality 2D reads of approximately 5500 bp that were up to 98% correctly assigned to the species level. Synthetic metagenome MinION libraries generated using version 6 kit and chemistry yielded from 899 to 3497 2D reads with lengths averaging 5700 bp with up to 98% assignment accuracy at the species level. The observed community proportions for “equal” and “rare” synthetic libraries were close to the known proportions, deviating from 0.1% to 10% across all tests. For a 20-species mock community with staggered contributions, a sequencing run detected all but 3 species (each included at <0.05% of DNA in the total mixture), 91% of reads were assigned to the correct species, 93% of reads were assigned to the correct genus, and >99% of reads were assigned to the correct family. Conclusions: At the current level of output and sequence quality (just under 4 × 103 2D reads for a synthetic metagenome), MinION sequencing followed by Kraken or One Codex analysis has the potential to provide rapid and accurate metagenomic analysis where the consortium is comprised of a limited number of taxa. Important considerations noted in this study included: high sensitivity of the MinION platform to the quality of input DNA, high variability of sequencing results across libraries and flow cells, and relatively small numbers of 2D reads per analysis limit. Together, these limited detection of very rare components of the microbial consortia, and would likely limit the utility of MinION for the sequencing of high-complexity metagenomic communities where thousands of taxa are expected. Furthermore, the limitations of the currently available data analysis tools suggest there is considerable room for improvement in the analytical approaches for the characterization of microbial communities using long reads. Nevertheless, the fact that the accurate taxonomic assignment of high-quality reads generated by MinION is approaching 99.5% and, in most cases, the inferred community structure mirrors the known proportions of a synthetic mixture warrants further exploration of practical application to environmental metagenomics as the platform continues to develop and improve. With further improvement in sequence throughput and error rate reduction, this platform shows great promise for precise real-time analysis of the composition and structure of more complex microbial communities.