Hap-seq: An Optimal Algorithm for Haplotype Phasing with Imputation Using Sequencing Data

Hap-seq: An Optimal Algorithm for Haplotype Phasing with Imputation Using Sequencing Data
复制标题

DOI:
10.1089/cmb.2012.0091
复制
发表时间:
2013-02-01
影响因子:
1.7
通讯作者:
Eskin, Eleazar
Eskin, Eleazar
中科院分区:
生物学4区
文献类型:
--
作者:
He, Dan;Han, Buhm;Eskin, Eleazar

文献摘要

被引文献

相似文献

单倍型的推断,或等位基因沿着每条染色体的顺序,是遗传学中的一个基本问题,对于许多分析都很重要,包括混合作图、通过血统鉴定同一性区域和插补。传统上,单倍型是从基因型数据推断得到的微阵列使用人口单倍型频率的信息推断从基因型个体的大样本或参考数据集,如HapMap。由于大的参考数据集的可用性,现代的单倍型定相方法沿着这些线是密切相关的插补方法。当应用于从测序研究中获得的数据时,获得单倍型的一种简单方法是首先从序列数据推断基因型,然后应用插补方法。然而,这种方法没有考虑到相同序列读段上的等位基因源自相同染色体。单倍型组装方法利用这种洞察力,并通过将读段分配给染色体来预测单倍型,其方式是使读段和预测的单倍型之间的冲突数量最小化。不幸的是,组装方法需要非常高的测序覆盖率,并且通常不能完全重建单倍型。在这项工作中,我们提出了一种新的方法,Hap-seq,它同时是一种插补和组装方法,使用似然框架将来自参考数据集的信息与来自读段的信息相结合。我们的方法应用动态编程算法来识别预测的单倍型,其最大化单倍型相对于参考数据集和单倍型相对于观察到的读数的联合似然。我们表明,我们的方法只需要低测序覆盖率,可以重建单倍型包含常见和罕见的等位基因,具有更高的准确性相比,国家的最先进的插补方法。
Inference of haplotypes, or the sequence of alleles along each chromosome, is a fundamental problem in genetics and is important for many analyses, including admixture mapping, identifying regions of identity by descent, and imputation. Traditionally, haplotypes are inferred from genotype data obtained from microarrays using information on population haplotype frequencies inferred from either a large sample of genotyped individuals or a reference dataset such as the HapMap. Since the availability of large reference datasets, modern approaches for haplotype phasing along these lines are closely related to imputation methods. When applied to data obtained from sequencing studies, a straightforward way to obtain haplotypes is to first infer genotypes from the sequence data and then apply an imputation method. However, this approach does not take into account that alleles on the same sequence read originate from the same chromosome. Haplotype assembly approaches take advantage of this insight and predict haplotypes by assigning the reads to chromosomes in such a way that minimizes the number of conflicts between the reads and the predicted haplotypes. Unfortunately, assembly approaches require very high sequencing coverage and are usually not able to fully reconstruct the haplotypes. In this work, we present a novel approach, Hap-seq, which is simultaneously an imputation and assembly method that combines information from a reference dataset with the information from the reads using a likelihood framework. Our method applies a dynamic programming algorithm to identify the predicted haplotype, which maximizes the joint likelihood of the haplotype with respect to the reference dataset and the haplotype with respect to the observed reads. We show that our method requires only low sequencing coverage and can reconstruct haplotypes containing both common and rare alleles with higher accuracy compared to the state-of-the-art imputation methods.