Genotype calling from next-generation sequencing data using haplotype information of reads.

Genotype calling from next-generation sequencing data using haplotype information of reads.
复制标题

使用读数的单倍型信息从下一代测序数据中进行基因型调用。

DOI:
10.1093/bioinformatics/bts047
复制
发表时间:
2012
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Zhang,Kui
Zhang,Kui
中科院分区:
--
文献类型:
--
作者:
Zhi,Degui;Wu,Jihua;Liu,Nianjun;Zhang,Kui

文献摘要

相似文献

动机:低覆盖度测序为全基因组测序提供了一种经济的策略。对一组个体进行测序时,由于测序覆盖率较低,基因型识别可能具有挑战性。基于连锁不平衡 (LD) 的基因分型调用的细化对于提高准确性至关重要。当前基于 LD 的方法使用单个潜在多态性位点 (PPS) 的读数计数或基因型可能性。跨越多个 PPS 的读取(跳跃读取)可以提供当前方法忽略的额外单倍型信息。 结果:在本文中,我们介绍了一种新的基于隐马尔可夫模型 (HMM) 的方法,该方法可以考虑跨相邻 PPS 的跳跃读取信息,并在 HapSeq 程序中实现它。我们的方法扩展了 Thunder 中的 HMM,并明确地将跳跃读取信息建模为以相邻 PPS 状态为条件的发射概率。我们的模拟结果表明,与Thunder相比,HapSeq将基因分型错误率降低了30%,从0.86%降低到0.60%。千人基因组计划的结果表明,HapSeq 将欧洲和非洲血统个体的基因分型错误率分别从 2.24% 和 2.76% 降低到 1.97% 和 2.50%,分别降低了 12% 和 9%。我们希望我们的计划能够提高大量正在进行和计划中的全基因组测序项目的基因分型质量。联系方式:dzhi@ms.soph.uab.edu; kzhang@ms.soph.uab.edu 可用性:HapSeq 软件包及其手册可在 www.ssg.uab.edu/hapseq/ 上找到和下载。补充信息:补充数据可在 Bioinformaticsonline 上获得。
Motivation:Low coverage sequencing provides an economic strategy for whole genome sequencing. When sequencing a set of individuals, genotype calling can be challenging due to low sequencing coverage. Linkage disequilibrium (LD) based refinement of genotyping calling is essential to improve the accuracy. Current LD-based methods use read counts or genotype likelihoods at individual potential polymorphic sites (PPSs). Reads that span multiple PPSs (jumping reads) can provide additional haplotype information overlooked by current methods.Results:In this article, we introduce a new Hidden Markov Model (HMM)-based method that can take into account jumping reads information across adjacent PPSs and implement it in the HapSeq program. Our method extends the HMM in Thunder and explicitly models jumping reads information as emission probabilities conditional on the states of adjacent PPSs. Our simulation results show that, compared to Thunder, HapSeq reduces the genotyping error rate by 30%, from 0.86% to 0.60%. The results from the 1000 Genomes Project show that HapSeq reduces the genotyping error rate by 12 and 9%, from 2.24% and 2.76% to 1.97% and 2.50% for individuals with European and African ancestry, respectively. We expect our program can improve genotyping qualities of the large number of ongoing and planned whole genome sequencing projects.Contact:dzhi@ms.soph.uab.edu; kzhang@ms.soph.uab.eduAvailability:The software package HapSeq and its manual can be found and downloaded at www.ssg.uab.edu/hapseq/.Supplementary information:Supplementary data are available atBioinformaticsonline.