Accounting for missing data in the estimation of contemporary genetic effective population size (Ne)

Accounting for missing data in the estimation of contemporary genetic effective population size (Ne)
复制标题

DOI:
10.1111/1755-0998.12049
复制
发表时间:
2013-03-01
影响因子:
7.7
通讯作者:
Ovenden, J. . R.
Ovenden, J. . R.
中科院分区:
生物学1区
文献类型:
--
作者:
Peel, D.;Waples, R. S.;Ovenden, J. . R.

文献摘要

被引文献

相似文献

理论模型经常应用于群体遗传数据集,而没有充分考虑缺失数据的影响。研究人员可以通过移除未能产生基因型别的个体和/或移除未能产生等位基因决定的基因座来处理缺失的数据,但尽管他们尽了最大努力,大多数数据集仍然包含一些缺失的数据。因此,实现的样本量在不同的位置是不同的,这给必须显式地考虑随机抽样误差的无偏方法带来了问题。计算当代有效人口数量(Ne)的一个常用解决方案是计算有效样本量,即未加权平均数或跨区域调和平均数。这是不理想的,因为它没有考虑到这样一个事实,即具有不同等位基因数量的基因座具有不同的信息含量。在这里,我们考虑当代有效种群数量(Ne)的遗传估计的这个问题。为了评估几种处理缺失数据的统计方法的偏差和精度,我们模拟了具有已知Ne和不同程度缺失数据的总体。在所有情况下,一种校正缺失数据的方法(固定-逆方差加权调和均值)在估计Ne的单样本和两样本(时间)方法中始终表现最好,并优于目前广泛使用的一些方法。这里采用的方法可能是调整其他群体遗传学方法的起点,这些方法包括每个座位样本量的分量。
Theoretical models are often applied to population genetic data sets without fully considering the effect of missing data. Researchers can deal with missing data by removing individuals that have failed to yield genotypes and/or by removing loci that have failed to yield allelic determinations, but despite their best efforts, most data sets still contain some missing data. As a consequence, realized sample size differs among loci, and this poses a problem for unbiased methods that must explicitly account for random sampling error. One commonly used solution for the calculation of contemporary effective population size (Ne) is to calculate the effective sample size as an unweighted mean or harmonic mean across loci. This is not ideal because it fails to account for the fact that loci with different numbers of alleles have different information content. Here we consider this problem for genetic estimators of contemporary effective population size (Ne). To evaluate bias and precision of several statistical approaches for dealing with missing data, we simulated populations with known Ne and various degrees of missing data. Across all scenarios, one method of correcting for missing data (fixed-inverse variance-weighted harmonic mean) consistently performed the best for both single-sample and two-sample (temporal) methods of estimating Ne and outperformed some methods currently in widespread use. The approach adopted here may be a starting point to adjust other population genetics methods that include per-locus sample size components.