Assessing statistical power of SNPs for population structure and conservation studies

Assessing statistical power of SNPs for population structure and conservation studies
复制标题

DOI:
10.1111/j.1755-0998.2008.02392.x
复制
发表时间:
2009-01-01
影响因子:
7.7
通讯作者:
Taylor, Barbara L.
Taylor, Barbara L.
中科院分区:
生物学1区
文献类型:
--
作者:
Morin, Phillip A.;Martien, Karen K.;Taylor, Barbara L.

文献摘要

被引文献

相似文献

一些人提出单核苷酸多态性 (SNP) 作为群体研究的新前沿,并且有几篇论文提出了报告 SNP 优点和局限性的理论和经验证据。然而,实际上,仍不清楚需要多少个 SNP 标记,或者这些标记的最佳特征应该是什么,才能获得足够的统计功效来检测不同水平的群体分化。我们使用一个假设的案例来说明设计群体遗传学项目的过程,并提供模拟结果,这些模拟结果解决了最大化统计能力以检测差异的几个问题,同时最大限度地减少了开发 SNP 的工作量。结果表明 (i) 虽然 30 个 SNP 应足以检测中等(F-ST = 0.01)水平的分化,但旨在检测人口统计独立性(例如 F-ST < 0.005)的研究可能需要 80 个或更多 SNP 和大样本量; (ii) 不同的SNP等位基因频率对功效的影响很小,因此,SNP的选择可以相对公正; (iii) 增加样本量对功效有很强的影响,因此当样本数已知时可以最小化基因座的数量,并且增加样本量几乎总是有益的; (iv) 通过在基因座内包含多个 SNP 并推断单倍型来提高功效,而不是尝试仅使用未连锁的 SNP。这还具有减少 SNP 确定工作的实际好处,并且可能影响是否在编码区域或非编码区域中寻找 SNP 的决定。
Single nucleotide polymorphisms (SNPs) have been proposed by some as the new frontier for population studies, and several papers have presented theoretical and empirical evidence reporting the advantages and limitations of SNPs. As a practical matter, however, it remains unclear how many SNP markers will be required or what the optimal characteristics of those markers should be in order to obtain sufficient statistical power to detect different levels of population differentiation. We use a hypothetical case to illustrate the process of designing a population genetics project, and present results from simulations that address several issues for maximizing statistical power to detect differentiation while minimizing the amount of effort in developing SNPs. Results indicate that (i) while 30 SNPs should be sufficient to detect moderate (F-ST = 0.01) levels of differentiation, studies aimed at detecting demographic independence (e.g. F-ST < 0.005) may require 80 or more SNPs and large sample sizes; (ii) different SNP allele frequencies have little affect on power, and thus, selection of SNPs can be relatively unbiased; (iii) increasing the sample size has a strong effect on power, so that the number of loci can be minimized when sample number is known, and increasing sample size is almost always beneficial; and (iv) power is increased by including multiple SNPs within loci and inferring haplotypes, rather than trying to use only unlinked SNPs. This also has the practical benefit of reducing the SNP ascertainment effort, and may influence the decision of whether to seek SNPs in coding or noncoding regions.