Bayesian haplotype inference for multiple linked single-nucleotide polymorphisms

Bayesian haplotype inference for multiple linked single-nucleotide polymorphisms
复制标题

DOI:
10.1086/338446
复制
发表时间:
2002-01-01
影响因子:
9.8
通讯作者:
Liu, JS
Liu, JS
中科院分区:
生物学1区
文献类型:
--
作者:
Niu, TH;Qin, ZHS;Liu, JS

文献摘要

被引文献

相似文献

由于单核苷酸多态性(SNP)和常规单位分析的有限功率,因此,单倍型在复杂蛋白酶基因的映射中引起了越来越多的关注。已经表明,诸如Clark算法,期望最大化算法和基于共聚的基于共聚的迭代抽样算法等单倍型的推荐方法是分子 - 型型型方法的相当有效且经济的替代方法。为了应对现有算法的一些弱点,我们提出了一种新的蒙特卡洛方法。特别是,我们将整个单倍型首先分为较小的细分市场。然后,我们使用Gibbs采样器均构建每个段的部分单倍型并将所有段组装在一起。对于大量链接的SNP,我们的算法可以准确,快速地推断单倍型。通过使用各种各样的真实和模拟数据集,我们证明了贝叶斯算法的优势,并且我们表明,违反Hardy-Weinberg平衡的违反,与丢失数据的存在以及重组热点的发生是强大的。
Haplotypes have gained increasing attention in the mapping of complex-disease genes, because of the abundance of single-nucleotide polymorphisms (SNPs) and the limited power of conventional single-locus analyses. It has been shown that haplotype-inference methods such as Clark's algorithm, the expectation-maximization algorithm, and a coalescence-based iterative-sampling algorithm are fairly effective and economical alternatives to molecular-haplotyping methods. To contend with some weaknesses of the existing algorithms, we propose a new Monte Carlo approach. In particular, we first partition the whole haplotype into smaller segments. Then, we use the Gibbs sampler both to construct the partial haplotypes of each segment and to assemble all the segments together. Our algorithm can accurately and rapidly infer haplotypes for a large number of linked SNPs. By using a wide variety of real and simulated data sets, we demonstrate the advantages of our Bayesian algorithm, and we show that it is robust to the violation of Hardy-Weinberg equilibrium, to the presence of missing data, and to occurrences of recombination hotspots.