Inferring Coancestry in Population Samples in the Presence of Linkage Disequilibrium

Inferring Coancestry in Population Samples in the Presence of Linkage Disequilibrium
复制标题

DOI:
10.1534/genetics.111.137570
复制
发表时间:
2012-04-01
期刊:
影响因子:
3.3
通讯作者:
Thompson, E. A.
Thompson, E. A.
中科院分区:
生物学2区
文献类型:
--
作者:
Brown, M. D.;Glazner, C. G.;Thompson, E. A.

文献摘要

被引文献

相似文献

在系谱连锁研究和基于群体的关联研究中,人们对利用现代密集遗传标记数据在未知亲缘关系的个体中推断遗传身份片段(ibd)非常感兴趣,以提高定位影响复杂性状的基因的能力和分辨率。在这篇文章中,我们提出了一种ibd在一组染色体之间的隐马尔可夫模型(HMM),并描述了在一对个体的四条染色体之间推断ibd的方法和软件,使用阶段(单倍型)或非阶段(基因型)数据。该模型允许丢失数据和输入错误,但不模拟连锁不平衡(LD),因为拟合准确的LD模型需要从经过充分研究的群体中获得大量样本。然而,遗传变异仍然是一个主要的混淆因素,因为遗传变异本身就反映了种群水平上的同祖先。为了研究LD的影响,我们开发了一种新的模拟方法,以在不同的LD水平上为同一组标记生成真实的密集标记数据。使用这种方法,我们展示了LD对我们的HMM模型在估计四个染色体组之间和基因型对之间ibd片段的敏感性和特异性的影响的研究结果。我们表明,尽管没有纳入LD,我们的模型在检测小至10(6)bp (1 Mpb)的片段方面非常成功;我们还比较了使用LD模型估计ibd的fastIBD。
In both pedigree linkage studies and in population-based association studies there has been much interest in the use of modern dense genetic marker data to infer segments of gene identity by descent (ibd) among individuals not known to be related, to increase power and resolution in localizing genes affecting complex traits. In this article, we present a hidden Markov model (HMM) for ibd among a set of chromosomes and describe methods and software for inference of ibd among the four chromosomes of pairs of individuals, using either phased (haplotypic) or unphased (genotypic) data. The model allows for missing data and typing error, but does not model linkage disequilibrium (LD), because fitting an accurate LD model requires large samples from well-studied populations. However, LD remains a major confounding factor, since LD is itself a reflection of coancestry at the population level. To study the impact of LD, we have developed a novel simulation approach to generate realistic dense marker data for the same set of markers but at varying levels of LD. Using this approach, we present results of a study of the impact of LD on the sensitivity and specificity of our HMM model in estimating segments of ibd among sets of four chromosomes and between genotype pairs. We show that, despite not incorporating LD, our model has been quite successful in detecting segments as small as 10(6) bp (1 Mpb); we present also comparisons with fastIBD which uses an LD model in estimating ibd.