Correcting for ascertainment bias in the inference of population structure

Correcting for ascertainment bias in the inference of population structure
复制标题

DOI:
10.1093/bioinformatics/btn665
复制
发表时间:
2009-02-15
期刊:
影响因子:
5.8
通讯作者:
Foll, Matthieu
Foll, Matthieu
中科院分区:
生物学3区
文献类型:
--
作者:
Guillot, Gilles;Foll, Matthieu

文献摘要

被引文献

相似文献

背景:分子标记的确定过程相当于忽略了携带频率较低的等位基因的基因座。如果推理算法没有适当考虑,这可能会导致群体遗传学模型下的推理出现强烈偏差。尝试对这种审查过程进行建模以推断人口结构(即识别个体集群)会带来具有挑战性的数值困难。方法:这些困难与大都会-黑斯廷斯接受率中存在棘手的归一化常数有关。这可以通过称为单变量交换算法 (SVEA) 的马尔可夫链蒙特卡罗 (MCMC) 算法来解决。结果:我们展示了如何针对群体遗传学中广泛关注的一类聚类模型实现此通用解决方案,其中包括计算机程序 STRUCTURE、GENELAND 和 GESTE 的基础模型。我们还实现了一个简单示例所提出的方法,并表明它允许我们大幅减少偏差。
Background: The ascertainment process of molecular markers amounts to disregard loci carrying alleles with low frequencies. This can result in strong biases in inferences under population genetics models if not properly taken into account by the inference algorithm. Attempting to model this censoring process in view of making inference of population structure (i.e. identifying clusters of individuals) brings up challenging numerical difficulties.Method: These difficulties are related to the presence of intractable normalizing constants in Metropolis-Hastings acceptance ratios. This can be solved via an Markov chain Monte Carlo (MCMC) algorithm known as single variable exchange algorithm (SVEA).Result: We show how this general solution can be implemented for a class of clustering models of broad interest in population genetics that includes the models underlying the computer programs STRUCTURE, GENELAND and GESTE. We also implement the method proposed for a simple example and show that it allows us to reduce the bias substantially.