On the design and analysis of next-generation sequencing genotyping for a cohort with haplotype-informative reads.

On the design and analysis of next-generation sequencing genotyping for a cohort with haplotype-informative reads.
复制标题

DOI:
10.1016/j.ymeth.2015.01.016
复制
发表时间:
2015-06
期刊:
影响因子:
4.8
通讯作者:
Zhang, Kui
Zhang, Kui
中科院分区:
生物学3区
文献类型:
--
作者:
Zhi, Degui;Liu, Nianjun;Zhang, Kui

文献摘要

参考文献

被引文献

相似文献

新一代测序(NGS)技术可以为所有类型的遗传变异提供碱基对分辨率的遗传信息,在遗传学研究中得到越来越多的应用。然而,由于NGS技术和分析的复杂性及其相对较高的成本,研究人员在设计和分析方面都面临着实际挑战。这些挑战由于最近的方法发展而变得更加复杂,这使得在测序中使用单倍型信息成为可能。鉴于这些进展,我们进行了全面的模拟,以评估测序覆盖率、配对末端读段插入长度和样本量对NGS研究中基因型召唤和单倍型分期的影响。与以往的研究通常使用理想化的场景来梳理个体设计和分析决策的影响相比,我们使用了一个完整的分析管道,从读取映射和变异检测到基因型调用和单倍型相位,因此我们可以评估多种决策的联合效应,从而为研究者提供更现实的建议。与之前的研究一致,我们发现在reads中使用单倍型信息可以提高基因型召唤和单倍型相位的准确性,并且我们还发现配对末端reads的长短插入长度混合可能提供更高的准确性。然而,只有在变异检测接近完美的高覆盖率测序中,这种益处才很明显。最后,我们观察到基于ld的优化方法并不总是优于基于单位点的基因型调用方法。因此,我们应该选择适合测序覆盖率和样本量的分析方法,以便在测序reads中使用单倍型信息。
Next-generation sequencing (NGS) technologies, which can provide base-pair resolution genetic information for all types of genetic variations, are increasingly used in genetics research. However, due to the complex nature of NGS technologies and analytics and their relatively high cost, investigators face practical challenges for both design and analysis. These challenges are further complicated by recent methodological developments that make it possible to use haplotype information in sequencing reads. In light of these developments, we conducted comprehensive simulations to evaluate the effects of sequencing coverage, insert size of paired-end reads, and sample size on genotype calling and haplotype phasing in NGS studies. In contrast to previous studies that typically use idealized scenarios to tease out the effects of individual design and analytic decisions, we used a complete analytical pipeline from read mapping and variant detection to genotype calling and haplotype phasing so that we can assess the joint effects of multiple decisions and thus make more realistic recommendations to investigators. Consistent with previous studies, we found that the use of haplotype information in reads can improve the accuracy of genotype calling and haplotype phasing, and we also found that a mixture of short and long insert sizes of paired-end reads may offer even greater accuracy. However, this benefit is only clear in high coverage sequencing where variant detection is close to perfect. Finally, we observed that LD-based refinement methods do not always outperform single site based methods for genotype calling. Therefore, we should choose analytical methods that are appropriate to the sequencing coverage and sample size in order to use haplotype information in sequencing reads.
DOI: 10.1126/science.1224344
发表时间: 2012-10-12
期刊: Science (New York, N.Y.)
影响因子: --
作者:
Meyer M;Kircher M;Gansauge MT;Li H;Racimo F;Mallick S;Schraiber JG;Jay F;Prüfer K;de Filippo C;Sudmant PH;Alkan C;Fu Q;Do R;Rohland N;Tandon A;Siebauer M;Green RE;Bryc K;Briggs AW;Stenzel U;Dabney J;Shendure J;Kitzman J;Hammer MF;Shunkov MV;Derevianko AP;Patterson N;Andrés AM;Eichler EE;Slatkin M;Reich D;Kelso J;Pääbo S
通讯作者: Pääbo S
来自1,092个人基因组的遗传变异的综合图。
DOI: 10.1038/nature11632
发表时间: 2012-11-01
期刊: Nature
影响因子: 64.8
作者:
通讯作者: --
DOI: 10.1093/bioinformatics/btr509
发表时间: 2011-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, Heng
通讯作者: Li, Heng
使用下一代 DNA 测序数据进行变异发现和基因分型的框架。
DOI: 10.1038/ng.806
发表时间: 2011-05
期刊: Nature genetics
影响因子: 30.8
作者:
通讯作者: --
DOI: 10.1002/gepi.20533
发表时间: 2010-12
影响因子: 2.1
作者:
Li, Yun;Willer, Cristen J.;Ding, Jun;Scheet, Paul;Abecasis, Goncalo R.
通讯作者: Abecasis, Goncalo R.