Characterization of sequence-specific errors in various next-generation sequencing systems

Characterization of sequence-specific errors in various next-generation sequencing systems
复制标题

DOI:
10.1039/c5mb00750j
复制
发表时间:
2016-01-01
影响因子:
--
通讯作者:
Park, Joonhong
Park, Joonhong
中科院分区:
生物3区
文献类型:
--
作者:
Shin, Sunguk;Park, Joonhong

文献摘要

被引文献

相似文献

下一代测序(NGS)是一种流行的方法,用于评估微生物群落的分子多样性,而无需培养,用于鉴定种群中的多态性,以及用于比较基因组和转录组。然而,NGS系统的序列特异性错误(SSE)可导致基因组错误组装、微生物群落分析中多样性的高估以及错误的多态性发现。由于丰富的微生物生物多样性和含有频繁重复的基因组,SSE可能特别成问题。在这项研究中,使用马尔可夫链模型发现了所有流行的NGS系统的公共数据中的SSE,并确定了序列错误的热点。在非Illumina NGS系统(如GS FLX+)中,缺失错误通常发生在均聚物之前。取代错误通常与Illumina测序系统如HiSeq中的高GC含量和长G/C均聚物有关。在HiSeq中去除长G/C均聚物后,重叠群的平均长度和平均SNP质量增加。通过质量过滤从我们的模拟社区数据中选择性地去除SSE,并确定了对特定微生物的偏见。我们的研究结果为过滤低质量读数、纠正缺失错误、防止基因组错误组装以及准确评估微生物群落组成和多态性提供了科学依据。
Next-generation sequencing (NGS) is a popular method for assessing the molecular diversity of microbial communities without cultivation, for identifying polymorphisms in populations, and for comparing genomes and transcriptomes. However, sequence-specific errors (SSEs) by NGS systems can result in genome mis-assembly, overestimation of diversity in microbial community analyses, and false polymorphism discovery. SSEs can be particularly problematic due to rich microbial biodiversity and genomes containing frequent repeats. In this study, SSEs in public data from all popular NGS systems were discovered using a Markov chain model and hotspots for sequence errors were identified. Deletion errors were frequently preceded by homopolymers in non-Illumina NGS systems, such as GS FLX+. Substitution errors were often related to high GC contents and long G/C homopolymers in Illumina sequencing systems such as HiSeq. After removal of long G/C homopolymers in HiSeq, the average lengths of contigs and average SNP quality increased. SSEs were selectively removed from our mock community data by quality filtering, and a bias against specific microbes was identified. Our findings provide a scientific basis for filtering poor-quality reads, correcting deletion errors, preventing genome mis-assembly, and accurately assessing microbial community compositions and polymorphisms.