On the design and analysis of next-generation sequencing genotyping for a cohort with haplotype-informative reads.
On the design and analysis of next-generation sequencing genotyping for a cohort with haplotype-informative reads.
复制标题
DOI:
10.1016/j.ymeth.2015.01.016
复制
发表时间:
2015-06
期刊:
影响因子:
4.8
通讯作者:
Zhang, Kui
中科院分区:
文献类型:
--
作者:
Zhi, Degui;Liu, Nianjun;Zhang, Kui
关键词:
Next-generation sequencing (NGS) technologies, which can provide base-pair resolution genetic information for all types of genetic variations, are increasingly used in genetics research. However, due to the complex nature of NGS technologies and analytics and their relatively high cost, investigators face practical challenges for both design and analysis. These challenges are further complicated by recent methodological developments that make it possible to use haplotype information in sequencing reads. In light of these developments, we conducted comprehensive simulations to evaluate the effects of sequencing coverage, insert size of paired-end reads, and sample size on genotype calling and haplotype phasing in NGS studies. In contrast to previous studies that typically use idealized scenarios to tease out the effects of individual design and analytic decisions, we used a complete analytical pipeline from read mapping and variant detection to genotype calling and haplotype phasing so that we can assess the joint effects of multiple decisions and thus make more realistic recommendations to investigators. Consistent with previous studies, we found that the use of haplotype information in reads can improve the accuracy of genotype calling and haplotype phasing, and we also found that a mixture of short and long insert sizes of paired-end reads may offer even greater accuracy. However, this benefit is only clear in high coverage sequencing where variant detection is close to perfect. Finally, we observed that LD-based refinement methods do not always outperform single site based methods for genotype calling. Therefore, we should choose analytical methods that are appropriate to the sequencing coverage and sample size in order to use haplotype information in sequencing reads.
登录
查看更多内容
DOI:
10.1126/science.1224344
发表时间:
2012-10-12
期刊:
Science (New York, N.Y.)
影响因子:
--
作者:
Meyer M;Kircher M;Gansauge MT;Li H;Racimo F;Mallick S;Schraiber JG;Jay F;Prüfer K;de Filippo C;Sudmant PH;Alkan C;Fu Q;Do R;Rohland N;Tandon A;Siebauer M;Green RE;Bryc K;Briggs AW;Stenzel U;Dabney J;Shendure J;Kitzman J;Hammer MF;Shunkov MV;Derevianko AP;Patterson N;Andrés AM;Eichler EE;Slatkin M;Reich D;Kelso J;Pääbo S
通讯作者:
Pääbo S
影响因子:
64.8
作者:
通讯作者:
--
影响因子:
5.8
作者:
Li, Heng
通讯作者:
Li, Heng
影响因子:
30.8
作者:
通讯作者:
--
影响因子:
2.1
作者:
Li, Yun;Willer, Cristen J.;Ding, Jun;Scheet, Paul;Abecasis, Goncalo R.
通讯作者:
Abecasis, Goncalo R.