Genotype Error Detection Using Hidden Markov Models of Haplotype Diversity

Genotype Error Detection Using Hidden Markov Models of Haplotype Diversity
复制标题

利用单倍型多样性的隐马尔可夫模型检测基因型错误

DOI:
10.1089/cmb.2007.0133
复制
发表时间:
2008-11-01
影响因子:
1.7
通讯作者:
Pasaniuc, Bogdan
Pasaniuc, Bogdan
中科院分区:
生物学4区
文献类型:
--
作者:
Kennedy, Justin;Mandoiu, Ion;Pasaniuc, Bogdan

文献摘要

被引文献

相似文献

基因分型错误的存在可能会使连锁和疾病关联的统计测试无效,特别是对于基于单倍型分析的方法。贝克尔等人。最近提出了一种简单的似然比方法来检测三基因型数据中的错误。在此方法下,当测试中的SNP基因型被删除时,如果与原始三基因型数据相关的可能性增加超过用户选择的阈值的乘法因子,则SNP基因型被标记为潜在错误。在本文中,我们使用似然比检验方法与似然函数相结合,给出了改进的错误检测方法,该方法可以基于所研究群体中单倍型多样性的隐马尔可夫模型进行有效计算。模拟和真实数据集上的实验结果表明,所提出的方法具有高度可扩展的运行时间,并且与以前的方法相比,检测精度显着提高。
The presence of genotyping errors can invalidate statistical tests for linkage and disease association, particularly for methods based on haplotype analysis. Becker et al. have recently proposed a simple likelihood ratio approach for detecting errors in trio genotype data. Under this approach, a SNP genotype is flagged as a potential error if the likelihood associated with the original trio genotype data increases by a multiplicative factor exceeding a user selected threshold when the SNP genotype under test is deleted. In this article we give improved error detection methods using the likelihood ratio test approach in conjunction with likelihood functions that can be efficiently computed based on a Hidden Markov Model of haplotype diversity in the population under study. Experimental results on both simulated and real datasets show that proposed methods have highly scalable running time and achieve significantly improved detection accuracy compared to previous methods.