Combining Markers into Haplotypes Can Improve Population Structure Inference

Combining Markers into Haplotypes Can Improve Population Structure Inference
复制标题

DOI:
10.1534/genetics.111.131136
复制
发表时间:
2012-01-01
期刊:
影响因子:
3.3
通讯作者:
Jakobsson, Mattias
Jakobsson, Mattias
中科院分区:
生物学2区
文献类型:
--
作者:
Gattepaille, Lucie M.;Jakobsson, Mattias

文献摘要

被引文献

相似文献

高通量基因分型和测序技术可以为大量个体生成密集的遗传标记集。对于大多数物种,这些数据将包含连锁不平衡(LD)中的许多标记。为了利用这些数据进行人口结构推断,我们研究了使用单倍型构建的等位基因在单核苷酸多态性(SNP)相结合。我们介绍了一个统计数据来自信息论,分配(GIA),它量化的额外信息分配个人的人口使用单倍型数据相比,单独使用单个位点的信息增益。使用一个双位点双等位基因模型,我们表明,在连锁平衡的标记组合成单倍型总是会导致非阳性GIA,这表明,结合两个标记是不利于祖先推断。然而,对于LD中的基因座,GIA通常是阳性的,这表明可以通过将标记组合成单倍型来改善分配。使用GIA作为一个标准结合标记成单倍型,我们证明了模拟数据的一个显着改善分配个人的候选人群。对于我们调查的许多情况,使用单倍型数据,错误分配减少了26%至97%。对于来自法国和德国个体的经验数据,例如,使用单倍型可以将错误分配的个体减少73%。我们的研究结果可以用于具有挑战性的人口结构和分配问题,特别是对于大规模人口基因组数据可用的研究。
High-throughput genotyping and sequencing technologies can generate dense sets of genetic markers for large numbers of individuals. For most species, these data will contain many markers in linkage disequilibrium (LD). To utilize such data for population structure inference, we investigate the use of haplotypes constructed by combining the alleles at single-nucleotide polymorphisms (SNPs). We introduce a statistic derived from information theory, the gain of informativeness for assignment (GIA), which quantifies the additional information for assigning individuals to populations using haplotype data compared to using individual loci separately. Using a two-loci-two-allele model, we demonstrate that combining markers in linkage equilibrium into haplotypes always leads to non-positive GIA, suggesting that combining the two markers is not advantageous for ancestry inference. However, for loci in LD, GIA is often positive, suggesting that assignment can be improved by combining markers into haplotypes. Using GIA as a criterion for combining markers into haplotypes, we demonstrate for simulated data a significant improvement of assigning individuals to candidate populations. For the many cases that we investigate, incorrect assignment was reduced between 26% and 97% using haplotype data. For empirical data from French and German individuals, the incorrectly assigned individuals can, for example, be decreased by 73% using haplotypes. Our results can be useful for challenging population structure and assignment problems, in particular for studies where large-scale population-genomic data are available.