Ultra-deep, long-read nanopore sequencing of mock microbial community standards

Ultra-deep, long-read nanopore sequencing of mock microbial community standards
复制标题

DOI:
10.1093/gigascience/giz043
复制
发表时间:
2019-05-01
期刊:
影响因子:
9.2
通讯作者:
Loman, Nicholas J.
Loman, Nicholas J.
中科院分区:
生物学2区
文献类型:
--
作者:
Nicholls, Samuel M.;Quick, Joshua C.;Loman, Nicholas J.

文献摘要

被引文献

相似文献

背景资料:长测序读数信息丰富:有助于从头组装和参考作图,因此对微生物群落的研究具有巨大潜力。然而,用于分析长读取宏基因组数据的最佳方法是未知的。此外,生物信息学工具的严格评估受到缺乏来自具有已知组成的经验证样品的长读数据的阻碍。调查结果:我们用Oxford Nanopore GridION和PromethION对含有10种微生物物种(ZymoBIOMICS Microbial Community Standards)的2个市售模拟群落进行了测序。两个社区和10个单独的物种分离物也用Illumina技术测序。对于均匀分布和对数分布的群落,我们分别从2个GridION流动池生成14和16千兆碱基对,从2个PromethION流动池生成150和153千兆碱基对。在4次测序运行中,读取长度N50范围在5.3和5.4个双链酶对之间。基础调用和相应的信号数据可用(总共4.2 TB)。与Illumina测序分离株的比对证明了预期丰度的预期微生物菌种,最低丰度菌种的检测限低于50个细胞(GridION)。宏基因组的从头组装恢复了长的连续序列,而不需要诸如分箱的预处理技术。结论:我们从一个定义明确的模拟社区中提出了超深,长读纳米孔数据集。这些数据集将有助于开发用于长读宏基因组学的生物信息学方法以及验证和比较当前实验室和软件管道。
Background: Long sequencing reads are information-rich: aiding de novo assembly and reference mapping, and consequently have great potential for the study of microbial communities. However, the best approaches for analysis of long-read metagenomic data are unknown. Additionally, rigorous evaluation of bioinformatics tools is hindered by a lack of long-read data from validated samples with known composition. Findings: We sequenced 2 commercially available mock communities containing 10 microbial species (ZymoBIOMICS Microbial Community Standards) with Oxford Nanopore GridION and PromethION. Both communities and the 10 individual species isolates were also sequenced with Illumina technology. We generated 14 and 16 gigabase pairs from 2 GridION flowcells and 150 and 153 gigabase pairs from 2 PromethION flowcells for the evenly distributed and log-distributed communities, respectively. Read length N50 ranged between 5.3 and 5.4 kilobase pairs over the 4 sequencing runs. Basecalls and corresponding signal data are made available (4.2 TB in total). Alignment to Illumina-sequenced isolates demonstrated the expected microbial species at anticipated abundances, with the limit of detection for the lowest abundance species below 50 cells (GridION). De novo assembly of metagenomes recovered long contiguous sequences without the need for pre-processing techniques such as binning. Conclusions: We present ultra-deep, long-read nanopore datasets from a well-defined mock community. These datasets will be useful for those developing bioinformatics methods for long-read metagenomics and for the validation and comparison of current laboratory and software pipelines.