Estimating DNA polymorphism from next generation sequencing data with high error rate by dual sequencing applications.

Estimating DNA polymorphism from next generation sequencing data with high error rate by dual sequencing applications.
复制标题

通过双测序应用从高错误率的下一代测序数据中估计 DNA 多态性。

DOI:
10.1186/1471-2164-14-535
复制
发表时间:
2013-08-07
期刊:
影响因子:
4.4
通讯作者:
Wu CI
Wu CI
中科院分区:
生物学2区
文献类型:
--
作者:
He Z;Li X;Ling S;Fu YX;Hungate E;Shi S;Wu CI

文献摘要

参考文献

被引文献

相似文献

由于下一代测序(NGS)数据的错误率高且误差在位点间的分布不均匀,从NGS数据中准确估计DNA多态性(θ)一直是一个挑战。通过计算机模拟,我们比较了两种数据获取方法——分别对每个二倍体个体进行测序和对合并样本进行测序。在目前的NGS错误率下,单独对每个个体进行测序几乎没有优势,除非每个个体的覆盖率很高(bbb20倍)。因此,我们提出了一种新的方法来估计θ从汇集的样本已经受到两个单独的DNA测序轮。由于两种测序应用的错误通常不重叠,因此可以从测序错误中分离出低频多态性。仿真结果表明,在错误率较高、θ较低的情况下,双应用方法仍然是可靠的。在自然种群的研究中,测序覆盖通常是适度的(每个个体~2X),对合并样本的双重应用方法应该是一个合理的选择。
As the error rate is high and the distribution of errors across sites is non-uniform in next generation sequencing (NGS) data, it has been a challenge to estimate DNA polymorphism (θ) accurately from NGS data. By computer simulations, we compare the two methods of data acquisition - sequencing each diploid individual separately and sequencing the pooled sample. Under the current NGS error rate, sequencing each individual separately offers little advantage unless the coverage per individual is high (>20X). We hence propose a new method for estimating θ from pooled samples that have been subjected to two separate rounds of DNA sequencing. Since errors from the two sequencing applications are usually non-overlapping, it is possible to separate low frequency polymorphisms from sequencing errors. Simulation results show that the dual applications method is reliable even when the error rate is high and θ is low. In studies of natural populations where the sequencing coverage is usually modest (~2X per individual), the dual applications method on pooled samples should be a reasonable choice.
DOI: 10.1126/science.1215040
发表时间: 2012-02-17
期刊: Science (New York, N.Y.)
影响因子: --
作者:
MacArthur DG;Balasubramanian S;Frankish A;Huang N;Morris J;Walter K;Jostins L;Habegger L;Pickrell JK;Montgomery SB;Albers CA;Zhang ZD;Conrad DF;Lunter G;Zheng H;Ayub Q;DePristo MA;Banks E;Hu M;Handsaker RE;Rosenfeld JA;Fromer M;Jin M;Mu XJ;Khurana E;Ye K;Kay M;Saunders GI;Suner MM;Hunt T;Barnes IH;Amid C;Carvalho-Silva DR;Bignell AH;Snow C;Yngvadottir B;Bumpstead S;Cooper DN;Xue Y;Romero IG;1000 Genomes Project Consortium;Wang J;Li Y;Gibbs RA;McCarroll SA;Dermitzakis ET;Pritchard JK;Barrett JC;Harrow J;Hurles ME;Gerstein MB;Tyler-Smith C
通讯作者: Tyler-Smith C
DOI: 10.1093/bioinformatics/btr509
发表时间: 2011-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, Heng
通讯作者: Li, Heng
DOI: 10.1534/genetics.107.080630
发表时间: 2009-01-01
期刊: GENETICS
影响因子: 3.3
作者:
Jiang, Rong;Tavare, Simon;Marjoram, Paul
通讯作者: Marjoram, Paul
用于检测正选择下搭便车的复合测试。
DOI: 10.1093/molbev/msm119
发表时间: 2007-08-01
影响因子: 10.7
作者:
Zeng, Kai;Shi, Suhua;Wut, Chung-I
通讯作者: Wut, Chung-I
DOI: 10.2144/000113962
发表时间: 2012-12
期刊: BioTechniques
影响因子: 2.7
作者:
Coupland P;Chandra T;Quail M;Reik W;Swerdlow H
通讯作者: Swerdlow H