Usefulness of single nucleotide polymorphism data for estimating population parameters.

Usefulness of single nucleotide polymorphism data for estimating population parameters.
复制标题

单核苷酸多态性数据对于估计群体参数的有用性。

DOI:
10.1093/genetics/156.1.439
复制
发表时间:
2000
期刊:
影响因子:
3.3
通讯作者:
Felsenstein,J
Felsenstein,J
中科院分区:
生物学2区
文献类型:
--
作者:
Kuhner,MK;Beerli,P;Yamato,J;Felsenstein,J

文献摘要

被引文献

相似文献

只要知道单核苷酸多态性(SNP)的确定方式,就可以利用最大似然法进行参数估计,从而构建合适的似然公式。我们提出了几种抽样方法的这种可能性。作为对这些方法的测试,我们考虑使用 SNP 来估计参数 θ = 4Neμ(有效种群大小和每个位点突变率的缩放乘积),该参数与重建谱系的分支长度有关。对于无限量的数据,使用 SNP 数据的 ML 模型预计会产生一致的 θ 估计值。对于有限数量的数据,当 θ 高时,估计值是准确的,但当 θ 低时,估计值往往会向上偏差。如果存在重组且分析中不允许出现重组,则结果会另外偏向上,但可以通过将重组纳入分析来消除这种影响。被定义为实际样本中多态性位点的 SNP(样本 SNP)对于 θ 的估计比由同一群体中选择的一组(组 SNP)中的多态性定义的 SNP 更准确。将面板 SNP 错误地表述为样本 SNP 会导致 θ 的最大似然估计出现较大误差。收集 SNP 的研究人员应收集并保存有关确定方法的信息,以便准确分析数据。
Single nucleotide polymorphism (SNP) data can be used for parameter estimation via maximum likelihood methods as long as the way in which the SNPs were determined is known, so that an appropriate likelihood formula can be constructed. We present such likelihoods for several sampling methods. As a test of these approaches, we consider use of SNPs to estimate the parameter Θ = 4Neμ (the scaled product of effective population size and per-site mutation rate), which is related to the branch lengths of the reconstructed genealogy. With infinite amounts of data, ML models using SNP data are expected to produce consistent estimates of Θ. With finite amounts of data the estimates are accurate when Θ is high, but tend to be biased upward when Θ is low. If recombination is present and not allowed for in the analysis, the results are additionally biased upward, but this effect can be removed by incorporating recombination into the analysis. SNPs defined as sites that are polymorphic in the actual sample under consideration (sample SNPs) are somewhat more accurate for estimation of Θ than SNPs defined by their polymorphism in a panel chosen from the same population (panel SNPs). Misrepresenting panel SNPs as sample SNPs leads to large errors in the maximum likelihood estimate of Θ. Researchers collecting SNPs should collect and preserve information about the method of ascertainment so that the data can be accurately analyzed.