Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.

Substantial biases in ultra-short read data sets from high-throughput DNA sequencing.
复制标题

从高通量DNA测序中的超短偏差读取数据集中的实质性偏见。

DOI:
10.1093/nar/gkn425
复制
发表时间:
2008-09
影响因子:
14.9
通讯作者:
Himmelbauer, Heinz
Himmelbauer, Heinz
中科院分区:
生物学2区
文献类型:
--
作者:
Dohm, Juliane C.;Lottaz, Claudio;Borodina, Tatiana;Himmelbauer, Heinz

文献摘要

参考文献

被引文献

相似文献

新颖的测序技术允许快速生成大型序列数据集。这些技术可能会彻底改变遗传学和生物医学研究,但有必要对超短读取输出进行彻底的表征。我们生成并分析了两个 Illumina 1G 超短读长数据集,即来自 Beta vulgaris 基因组克隆的 280 万个 27mer 读长和来自 Helicobacter acinonychis 基因组的 1230 万个 36mer 读长。我们发现错误率范围从读取开始时的 0.3% 到读取结束时的 3.8%。错误的碱基识别经常出现在碱基 G 之前。碱基取代错误频率变化为 10 至 11 倍,其中 A > C 颠换是最常见的置换错误,C > G 颠换是最不常见的置换错误。单碱基的插入和删除发生率非常低。在模拟重测序时,我们发现 20 倍的测序覆盖度足以通过正确读取来补偿错误。测序区域的read覆盖率存在偏差;在 GC 含量升高的区间内发现了最高的读数密度。高 Solexa 质量分数表示过于乐观,低分数则低估了数据质量。我们的结果显示了不同类型的偏见以及检测它们的方法。这种偏差会对 Solexa 数据的使用和解释、从头测序、重测序、单核苷酸多态性和 DNA 甲基化位点的识别以及转录组分析产生影响。
Novel sequencing technologies permit the rapid production of large sequence data sets. These technologies are likely to revolutionize genetics and biomedical research, but a thorough characterization of the ultra-short read output is necessary. We generated and analyzed two Illumina 1G ultra-short read data sets, i.e. 2.8 million 27mer reads from a Beta vulgaris genomic clone and 12.3 million 36mers from the Helicobacter acinonychis genome. We found that error rates range from 0.3% at the beginning of reads to 3.8% at the end of reads. Wrong base calls are frequently preceded by base G. Base substitution error frequencies vary by 10- to 11-fold, with A > C transversion being among the most frequent and C > G transversions among the least frequent substitution errors. Insertions and deletions of single bases occur at very low rates. When simulating re-sequencing we found a 20-fold sequencing coverage to be sufficient to compensate errors by correct reads. The read coverage of the sequenced regions is biased; the highest read density was found in intervals with elevated GC content. High Solexa quality scores are over-optimistic and low scores underestimate the data quality. Our results show different types of biases and ways to detect them. Such biases have implications on the use and interpretation of Solexa data, for de novo sequencing, re-sequencing, the identification of single nucleotide polymorphisms and DNA methylation sites, as well as for transcriptome analysis.
DOI: 10.1186/gb-2007-8-7-r143
发表时间: 2007
期刊: Genome biology
影响因子: 12.3
作者:
Huse SM;Huber JA;Morrison HG;Sogin ML;Welch DM
通讯作者: Welch DM
DOI: 10.1126/science.1137325
发表时间: 2007-06-08
期刊: SCIENCE
影响因子: 56.9
作者:
Kim, Jae Bum;Porreca, Gregory J.;Seidman, J. G.
通讯作者: Seidman, J. G.
DOI: 10.1371/journal.pgen.0020120
发表时间: 2006-07-01
期刊: PLOS GENETICS
影响因子: 4.5
作者:
Eppinger, Mark;Baar, Claudia;Schuster, Stephan C.
通讯作者: Schuster, Stephan C.
DOI: 10.1101/sqb.1986.051.01.032
发表时间: 1986-01-01
期刊: COLD SPRING HARBOR SYMPOSIA ON QUANTITATIVE BIOLOGY
影响因子: --
作者:
MULLIS, K;FALOONA, F;ERLICH, H
通讯作者: ERLICH, H
测序的Medicago truncatula使用454个生命科学技术表达了测序标签。
DOI: 10.1186/1471-2164-7-272
发表时间: 2006-10-24
期刊: BMC GENOMICS
影响因子: 4.4
作者:
Cheung, Foo;Haas, Brian J;Goldberg, Susanne M D;May, Gregory D;Xiao, Yongli;Town, Christopher D
通讯作者: Town, Christopher D