A Markov chain Monte Carlo approach for joint inference of population structure and inbreeding rates from multilocus genotype data

A Markov chain Monte Carlo approach for joint inference of population structure and inbreeding rates from multilocus genotype data
复制标题

DOI:
10.1534/genetics.107.072371
复制
发表时间:
2007-07-01
期刊:
影响因子:
3.3
通讯作者:
Bustamante, Carlos D.
Bustamante, Carlos D.
中科院分区:
生物学2区
文献类型:
--
作者:
Gao, Hong;Williamson, Scott;Bustamante, Carlos D.

文献摘要

被引文献

相似文献

非随机交配诱导基因座内和基因座间等位基因状态的相关性,这些相关性可用于理解自然群体的遗传结构(WRIGHT 1965)。对于许多物种来说,量化两种形式的非随机交配对常设遗传变异模式的贡献是相当有意义的:近亲繁殖(亲属间的交配)和群体亚结构(配子的有限分散)。在这里,我们扩展了流行的贝叶斯聚类方法结构(PRITCHARD等2000年)的近亲繁殖或自交率和人口的起源分类使用多位点遗传标记的同时推理。这是通过消除集群内的Hardy-Weinberg平衡的假设来实现的,而是根据近交或自交率计算预期的基因型频率。我们证明了这样的扩展的必要性,表明自交导致虚假信号的人口子结构,使用标准的结构算法的混合物的虚假信号的偏见。我们使用广泛的聚结模拟来衡量我们的方法的性能,并证明我们的方法可以纠正这种偏差。我们还应用我们的方法来了解野生相对的驯化水稻,普通野生稻,一个重要的部分自交草种的种群结构。使用在111个随机位点测序的n = 16个个体的样本,我们发现存在两个亚群的强有力的证据,这与采样的地理位置密切相关,并且估计两个组的设定率与实验数据的估计值一致(s近似为0.48-0.70)。
Nonrandom mating induces correlations in allelic states within and among loci that can be exploited to understand the genetic structure of natural populations (WRIGHT 1965). For many species, it is of considerable interest to quantify the contribution of two forms of nonrandom mating to patterns of standing genetic variation: inbreeding (mating among relatives) and population substructure (limited dispersal of gametes). Here, we extend the popular Bayesian clustering approach STRUCTURE (PRITCHARD et al 2000) for simultaneous inference of inbreeding or selfing rates and population-of-origin classification using multilocus genetic markers. This is accomplished by eliminating the assumption of Hardy-Weinberg equilibrium within clusters and, instead, calculating expected genotype frequencies on the basis of inbreeding or selfing rates. We demonstrate the need for such an extension by showing that selfing leads to spurious signals Of population substructure using the standard STRUCTURE algorithm with a bias toward spurious signals of admixture. We gauge the performance of our method using extensive coalescent simulations and demonstrate that our approach can correct for this bias. We also apply our approach to understanding the population structure of the wild relative of domesticated rice, Oryza rufipogon, an important partially selfing grass species. Using a sample of n = 16 individuals sequenced at 1 1 1 random loci, we find strong evidence for existence of two subpopulations, which correlates well with geographic location of sampling, and estimate setting rates for both groups that are consistent with estimates from experimental data (s approximate to 0.48-0.70).