The effect of RAD allele dropout on the estimation of genetic variation within and between populations

The effect of RAD allele dropout on the estimation of genetic variation within and between populations
复制标题

DOI:
10.1111/mec.12089
复制
发表时间:
2013-06-01
期刊:
影响因子:
4.9
通讯作者:
Estoup, Arnaud
Estoup, Arnaud
中科院分区:
生物学1区
文献类型:
--
作者:
Gautier, Mathieu;Gharbi, Karim;Estoup, Arnaud

文献摘要

被引文献

相似文献

应用于减少代表性基因组的廉价短读段测序技术通过允许对大量个体和群体的大量单核苷酸多态性(SNP)进行基因分型,正在彻底改变遗传学研究,特别是群体遗传学分析。限制性位点相关DNA(RAD)测序是基于限制性位点侧翼的基因组区域的表征的最新技术。其潜在的缺点之一是限制性位点内多态性的存在,这使得不可能观察到相关的SNP等位基因(即等位基因缺失,ADO)。为了研究ADO对RAD标记估计的遗传变异的影响,我们首先从数学上推导出ADO对等位基因频率的影响的测量值,作为单个群体内不同参数的函数。然后,我们使用RAD数据集模拟使用聚结模型,调查ADO引起的偏差的大小估计的预期杂合性和FST下的一个简单的人口模型,两个群体之间的分歧。我们发现,ADO倾向于高估种群内和种群间的遗传变异。假设每个核苷酸的突变率在10-9和10-8之间,这种偏差对于大多数研究的发散时间和有效群体大小的组合来说仍然很低,除了大的有效群体大小。例如,通过滑动窗口分析对多个SNP的FST值进行平均,并不能纠正ADO偏差。我们简要地讨论了可能的解决方案,以过滤最有问题的情况下,ADO使用读取覆盖率检测标记大量过剩的无效等位基因。
Inexpensive short-read sequencing technologies applied to reduced representation genomes is revolutionizing genetic research, especially population genetics analysis, by allowing the genotyping of massive numbers of single-nucleotide polymorphisms (SNP) for large numbers of individuals and populations. Restriction site-associated DNA (RAD) sequencing is a recent technique based on the characterization of genomic regions flanking restriction sites. One of its potential drawbacks is the presence of polymorphism within the restriction site, which makes it impossible to observe the associated SNP allele (i.e. allele dropout, ADO). To investigate the effect of ADO on genetic variation estimated from RAD markers, we first mathematically derived measures of the effect of ADO on allele frequencies as a function of different parameters within a single population. We then used RAD data sets simulated using a coalescence model to investigate the magnitude of biases induced by ADO on the estimation of expected heterozygosity and FST under a simple demographic model of divergence between two populations. We found that ADO tends to overestimate genetic variation both within and between populations. Assuming a mutation rate per nucleotide between 10-9 and 10-8, this bias remained low for most studied combinations of divergence time and effective population size, except for large effective population sizes. Averaging FST values over multiple SNPs, for example, by sliding window analysis, did not correct ADO biases. We briefly discuss possible solutions to filter the most problematic cases of ADO using read coverage to detect markers with a large excess of null alleles.