A hidden Markov model for haplotype inference for present-absent data of clustered genes using identified haplotypes and haplotype patterns.

A hidden Markov model for haplotype inference for present-absent data of clustered genes using identified haplotypes and haplotype patterns.
复制标题

DOI:
10.3389/fgene.2014.00267
复制
发表时间:
2014
影响因子:
3.7
通讯作者:
Zhang K
Zhang K
中科院分区:
生物学3区
文献类型:
--
作者:
Wu J;Chen GB;Zhi D;Liu N;Zhang K

文献摘要

参考文献

相似文献

大多数杀伤细胞免疫球蛋白样受体(KIR)基因检测为存在或不存在使用位点特异性基因分型技术。由于特定KIR基因的确切拷贝数(一个或两个)是未知的,因此存在不确定性。因此,这些基因的单倍型推断变得更具挑战性,由于这样大部分的缺失信息。同时,许多单倍型和部分单倍型模式已被确定,由于这些聚集的基因之间的紧密连锁不平衡(LD),因此可以合并,以方便单倍型推断。在本文中,我们开发了一种基于隐马尔可夫模型(HMM)的方法,该方法可以结合识别的单倍型或部分单倍型模式,用于从聚类基因的存在-不存在数据(例如,KIR基因)。我们比较了它的性能与期望最大化(EM)为基础的方法,以前开发的单倍型分配和单倍型频率估计,通过广泛的模拟KIR基因。仿真结果表明,当单倍型中含有不正确的单倍型和/或单倍型频率的标准差较小时,基于HMM的新方法优于以往的方法。我们还比较了我们的方法的性能与两种方法,不使用以前确定的单倍型和单倍型模式,包括基于EM的方法,HPALORE,和基于HMM的方法,MaCH。我们的模拟结果表明,将识别的单倍型和部分单倍型模式,可以提高单倍型推断的准确性。新的软件包HaploHMM可在http://www.soph.uab.edu/ssg/files/People/KZhang/HaploHMM/haplohmm-index.html下载。
The majority of killer cell immunoglobin-like receptor (KIR) genes are detected as either present or absent using locus-specific genotyping technology. Ambiguity arises from the presence of a specific KIR gene since the exact copy number (one or two) of that gene is unknown. Therefore, haplotype inference for these genes is becoming more challenging due to such large portion of missing information. Meantime, many haplotypes and partial haplotype patterns have been previously identified due to tight linkage disequilibrium (LD) among these clustered genes thus can be incorporated to facilitate haplotype inference. In this paper, we developed a hidden Markov model (HMM) based method that can incorporate identified haplotypes or partial haplotype patterns for haplotype inference from present-absent data of clustered genes (e.g., KIR genes). We compared its performance with an expectation maximization (EM) based method previously developed in terms of haplotype assignments and haplotype frequency estimation through extensive simulations for KIR genes. The simulation results showed that the new HMM based method outperformed the previous method when some incorrect haplotypes were included as identified haplotypes and/or the standard deviation of haplotype frequencies were small. We also compared the performance of our method with two methods that do not use previously identified haplotypes and haplotype patterns, including an EM based method, HPALORE, and a HMM based method, MaCH. Our simulation results showed that the incorporation of identified haplotypes and partial haplotype patterns can improve accuracy for haplotype inference. The new software package HaploHMM is available and can be downloaded at http://www.soph.uab.edu/ssg/files/People/KZhang/HaploHMM/haplohmm-index.html.
DOI: 10.1038/ng2088
发表时间: 2007-07-01
期刊: NATURE GENETICS
影响因子: 30.8
作者:
Marchini, Jonathan;Howie, Bryan;Donnelly, Peter
通讯作者: Donnelly, Peter
DOI: 10.1093/bioinformatics/btq157
发表时间: 2010-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Su, Shu-Yi;Asher, Julian E.;Coin, Lachlan J. M.
通讯作者: Coin, Lachlan J. M.
DOI: 10.1016/s0198-8859(03)00067-3
发表时间: 2003-06-01
期刊: HUMAN IMMUNOLOGY
影响因子: 2.7
作者:
Marsh, SGE;Parham, P;Wain, H
通讯作者: Wain, H
DOI: 10.1002/gepi.20533
发表时间: 2010-12
影响因子: 2.1
作者:
Li, Yun;Willer, Cristen J.;Ding, Jun;Scheet, Paul;Abecasis, Goncalo R.
通讯作者: Abecasis, Goncalo R.
DOI: 10.1534/g3.111.000174
发表时间: 2011-06
期刊: G3 (Bethesda, Md.)
影响因子: --
作者:
Kato M;Yoon S;Hosono N;Leotta A;Sebat J;Tsunoda T;Zhang MQ
通讯作者: Zhang MQ