Pseudo-Sanger sequencing: massively parallel production of long and near error-free reads using NGS technology.

Pseudo-Sanger sequencing: massively parallel production of long and near error-free reads using NGS technology.
复制标题

伪桑格测序:使用 NGS 技术大规模并行生成长且几乎无错误的读数

DOI:
10.1186/1471-2164-14-711
复制
发表时间:
2013-10-17
期刊:
影响因子:
4.4
通讯作者:
Wu CI
Wu CI
中科院分区:
生物学2区
文献类型:
--
作者:
Ruan J;Jiang L;Chong Z;Gong Q;Li H;Li C;Tao Y;Zheng C;Zhai W;Turissini D;Cannon CH;Lu X;Wu CI

文献摘要

参考文献

被引文献

相似文献

通常,下一代测序(NGS)技术具有超高通量的特性,但与传统的Sanger测序相比,其读取长度明显短。对端NGS可以在计算上延长读取长度,但由于其固有的间隙,给实际应用带来了很多不便。既然Illumina配对端测序能够从600 bp甚至800 bp的DNA片段中读取两端,如何填补配对端之间的空白以产生准确的长读取是一个有趣但具有挑战性的问题。结果我们开发了一种新技术,称为伪桑格(PS)测序。它试图填补配对末端之间的空白,并可以产生接近无错误的序列,相当于传统的桑格读取长度,但具有下一代测序的高通量。PS方法的主要新颖之处在于间隙填充是基于两端有重叠的对端reads的局部组装。因此,我们能够正确地填补重复基因组区域的空白。PS测序从NGS平台的短片段开始,使用一系列插入大小逐步减少的成对端文库。引入了一种计算方法,将这些特殊的对端reads转化为长度与插入尺寸最大的序列相对应的长且接近无错误的PS序列。与未转化的reads相比,PS结构具有间隙填充、误差校正和杂合子耐受性3个优点。PS构建的众多应用之一是从头基因组组装,我们在本研究中进行了测试。来自果蝇非等基因菌株的PS reads的组装产生了190 kb的N50序列,比现有的从头组装方法改进了5倍,比来自454测序的长reads的组装优势3倍。结论sour方法可从NGS对端测序中获得接近无错误的长reads。我们证明了从头组装可以从这些Sanger-like reads中获益良多。此外,长reads的特性可以应用于结构变异检测和宏基因组学等应用。
BackgroundUsually, next generation sequencing (NGS) technology has the property of ultra-high throughput but the read length is remarkably short compared to conventional Sanger sequencing. Paired-end NGS could computationally extend the read length but with a lot of practical inconvenience because of the inherent gaps. Now that Illumina paired-end sequencing has the ability of read both ends from 600 bp or even 800 bp DNA fragments, how to fill in the gaps between paired ends to produce accurate long reads is intriguing but challenging.ResultsWe have developed a new technology, referred to as pseudo-Sanger (PS) sequencing. It tries to fill in the gaps between paired ends and could generate near error-free sequences equivalent to the conventional Sanger reads in length but with the high throughput of the Next Generation Sequencing. The major novelty of PS method lies on that the gap filling is based on local assembly of paired-end reads which have overlaps with at either end. Thus, we are able to fill in the gaps in repetitive genomic region correctly. The PS sequencing starts with short reads from NGS platforms, using a series of paired-end libraries of stepwise decreasing insert sizes. A computational method is introduced to transform these special paired-end reads into long and near error-free PS sequences, which correspond in length to those with the largest insert sizes. The PS construction has 3 advantages over untransformed reads: gap filling, error correction and heterozygote tolerance. Among the many applications of the PS construction is de novo genome assembly, which we tested in this study. Assembly of PS reads from a non-isogenic strain ofDrosophila melanogasteryields an N50 contig of 190 kb, a 5 fold improvement over the existing de novo assembly methods and a 3 fold advantage over the assembly of long reads from 454 sequencing.ConclusionsOur method generated near error-free long reads from NGS paired-end sequencing. We demonstrated that de novo assembly could benefit a lot from these Sanger-like reads. Besides, the characteristic of the long reads could be applied to such applications as structural variations detection and metagenomics.
DOI: 10.1038/nmeth.1416
发表时间: 2010-02
期刊: NATURE METHODS
影响因子: 48
作者:
Hiatt, Joseph B.;Patwardhan, Rupali P.;Turner, Emily H.;Lee, Choli;Shendure, Jay
通讯作者: Shendure, Jay
DOI: 10.1101/gr.097261.109
发表时间: 2010-02-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Li, Ruiqiang;Zhu, Hongmei;Wang, Jun
通讯作者: Wang, Jun
种群基因组学:果蝇果蝇中多态性和差异的全基因组分析。
DOI: 10.1371/journal.pbio.0050310
发表时间: 2007-11-06
期刊: PLOS BIOLOGY
影响因子: 9.8
作者:
Begun, David J.;Holloway, Alisha K.;Stevens, Kristian;Hillier, LaDeana W.;Poh, Yu-Ping;Hahn, Matthew W.;Nista, Phillip M.;Jones, Corbin D.;Kern, Andrew D.;Dewey, Colin N.;Pachter, Lior;Myers, Eugene;Langley, Charles H.
通讯作者: Langley, Charles H.
DOI: 10.1101/gr.089532.108
发表时间: 2009-06-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Simpson, Jared T.;Wong, Kim;Birol, Inanc
通讯作者: Birol, Inanc
DOI: 10.1186/1471-2105-8-64
发表时间: 2007-02-26
期刊: BMC bioinformatics
影响因子: 3
作者:
通讯作者: --