Insight into biases and sequencing errors for amplicon sequencing with the Illumina MiSeq platform.

Insight into biases and sequencing errors for amplicon sequencing with the Illumina MiSeq platform.
复制标题

DOI:
10.1093/nar/gku1341
复制
发表时间:
2015-03-31
影响因子:
14.9
通讯作者:
Quince C
Quince C
中科院分区:
生物学2区
文献类型:
--
作者:
Schirmer M;Ijaz UZ;D'Amore R;Hall N;Sloan WT;Quince C

文献摘要

参考文献

被引文献

相似文献

Illumina 的 MiSeq 目前的读长长度高达 2 × 300 bp,具有高通量和低测序成本,正在成为全球最常用的测序平台之一。即使对于小型实验室来说,该平台也易于管理且价格合理。这使得目标基因测序、宏基因组学、小基因组测序和临床分子诊断等广泛应用能够快速周转。然而,人们对 Illumina 错误配置文件的了解仍然很少,因此程序并不是针对 Illumina 数据的特性而设计的。更好地了解错误模式对于序列分析至关重要,对于我们得出有效的结论也至关重要。研究群体样本中真实的遗传变异对于了解疾病、进化和起源至关重要。我们基于 16S rRNA 扩增子测序数据对 MiSeq 的错误模式进行了大规模研究。我们测试了用于扩增子测序的最先进的文库制备方法,结果表明文库制备方法和引物的选择是最重要的偏差来源,并导致不同的错误模式。此外,我们测试了各种纠错策略的效率,并确定质量修剪 (Sickle) 与纠错 (BayesHammer) 相结合,然后读取重叠 (PANDAseq) 是最成功的方法,平均降低了 93% 的替换错误率。
With read lengths of currently up to 2 × 300 bp, high throughput and low sequencing costs Illumina's MiSeq is becoming one of the most utilized sequencing platforms worldwide. The platform is manageable and affordable even for smaller labs. This enables quick turnaround on a broad range of applications such as targeted gene sequencing, metagenomics, small genome sequencing and clinical molecular diagnostics. However, Illumina error profiles are still poorly understood and programs are therefore not designed for the idiosyncrasies of Illumina data. A better knowledge of the error patterns is essential for sequence analysis and vital if we are to draw valid conclusions. Studying true genetic variation in a population sample is fundamental for understanding diseases, evolution and origin. We conducted a large study on the error patterns for the MiSeq based on 16S rRNA amplicon sequencing data. We tested state-of-the-art library preparation methods for amplicon sequencing and showed that the library preparation method and the choice of primers are the most significant sources of bias and cause distinct error patterns. Furthermore we tested the efficiency of various error correction strategies and identified quality trimming (Sickle) combined with error correction (BayesHammer) followed by read overlapping (PANDAseq) as the most successful approach, reducing substitution error rates on average by 93%.
DOI: 10.1111/1462-2920.12086
发表时间: 2013-06
影响因子: 5.1
作者:
Shakya M;Quince C;Campbell JH;Yang ZK;Schadt CW;Podar M
通讯作者: Podar M
DOI: 10.1186/gb-2007-8-7-r143
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Huse SM;Huber JA;Morrison HG;Sogin ML;Welch DM
通讯作者: Welch DM
DOI: 10.1093/nar/gkr344
发表时间: 2011-07
影响因子: 14.9
作者:
Nakamura K;Oshima T;Morimoto T;Ikeda S;Yoshikawa H;Shiwa Y;Ishikawa S;Linak MC;Hirai A;Takahashi H;Altaf-Ul-Amin M;Ogasawara N;Kanaya S
通讯作者: Kanaya S
DOI: 10.1186/gb-2009-10-8-r83
发表时间: 2009
期刊: Genome biology
影响因子: 12.3
作者:
Kircher M;Stenzel U;Kelso J
通讯作者: Kelso J
DOI: 10.1093/bioinformatics/btt593
发表时间: 2014-03-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Zhang J;Kobert K;Flouri T;Stamatakis A
通讯作者: Stamatakis A