An MCMC algorithm for haplotype assembly from whole-genome sequence data

An MCMC algorithm for haplotype assembly from whole-genome sequence data
复制标题

DOI:
10.1101/gr.077065.108
复制
发表时间:
2008-08-01
期刊:
影响因子:
7
通讯作者:
Bafna, Vineet
Bafna, Vineet
中科院分区:
生物学1区
文献类型:
--
作者:
Bansal, Vikas;Halpern, Aaron L.;Bafna, Vineet

文献摘要

被引文献

相似文献

与基因类型相比,关于单倍型(单个染色体上存在的等位基因的组合)的知识对全基因组关联研究和对人类进化史的推断要有用得多。单倍型通常是使用计算方法从群体基因数据中推断出来的。全基因组序列数据为构建个体跨越数百个千碱基的单倍型提供了一个很有前途的资源。在这篇文章中,我们提出了一种马尔可夫链蒙特卡罗(MCMC)算法,HASH(单人单倍型组装),用于从已映射到参考基因组组装的测序DNA片段中组装单倍型。马尔可夫链的转变是使用对从测序片段导出的图的最小割计算来生成的。我们已经应用我们的方法,使用最近测序的人类个体的全基因组猎枪序列数据来推断单倍型。高序列覆盖率和配对的存在导致了相当长的单倍型(N50长度类似于350kb)。基于测序片段与个体单倍型的比较,我们证明了使用HASH推断的该个体的单倍型比使用先前提出的贪婪启发式和简单的MCMC方法估计的单倍型要准确得多。使用HapMap项目中的单倍型,我们估计使用HASH推断的单倍型的切换错误率相当低,类似于1.1%。我们的马尔可夫链蒙特卡罗算法代表了单倍型组装的一般框架,可以应用于其他测序技术产生的序列数据。实现这些方法和阶段性个体单倍型的代码可以从http://www.cse.ucsd.edu/users/vibansal/HASH/.下载
In comparison to genotypes, knowledge about haplotypes ( the combination of alleles present on a single chromosome) is much more useful for whole-genome association studies and for making inferences about human evolutionary history. Haplotypes are typically inferred from population genotype data using computational methods. Whole-genome sequence data represent a promising resource for constructing haplotypes spanning hundreds of kilobases for an individual. In this article, we propose a Markov chain Monte Carlo (MCMC) algorithm, HASH ( haplotype assembly for single human), for assembling haplotypes from sequenced DNA fragments that have been mapped to a reference genome assembly. The transitions of the Markov chain are generated using min-cut computations on graphs derived from the sequenced fragments. We have applied our method to infer haplotypes using whole-genome shotgun sequence data from a recently sequenced human individual. The high sequence coverage and presence of mate pairs result in fairly long haplotypes (N50 length similar to 350 kb). Based on comparison of the sequenced fragments against the individual haplotypes, we demonstrate that the haplotypes for this individual inferred using HASH are significantly more accurate than the haplotypes estimated using a previously proposed greedy heuristic and a simple MCMC method. Using haplotypes from the HapMap project, we estimate the switch error rate of the haplotypes inferred using HASH to be quite low, similar to 1.1%. Our Markov chain Monte Carlo algorithm represents a general framework for haplotype assembly that can be applied to sequence data generated by other sequencing technologies. The code implementing the methods and the phased individual haplotypes can be downloaded from http://www.cse.ucsd.edu/users/vibansal/HASH/.