Assessing the power of informative subsets of loci for population assignment: standard methods are upwardly biased

Assessing the power of informative subsets of loci for population assignment: standard methods are upwardly biased
复制标题

DOI:
10.1111/j.1755-0998.2010.02846.x
复制
发表时间:
2010-07-01
影响因子:
7.7
通讯作者:
Anderson, E. C.
Anderson, E. C.
中科院分区:
生物学1区
文献类型:
--
作者:
Anderson, E. C.

文献摘要

被引文献

相似文献

众所周知,统计分类程序应使用与用于训练分类器的数据分开的数据进行评估。当所讨论的分类程序是使用一组遗传标记进行群体分配时,这一原则通常被忽视,所述遗传标记是根据其等位基因频率从大量候选标记中特别选择的。这一疏忽导致了所选的一组用于群体分配的标记物的预测准确性的系统性向上偏倚。三个广泛使用的软件程序选择标记信息的人口分配遭受这种偏见。通过一小组模拟记录了这种偏差的程度。当从低分化群体中筛选许多候选基因座时,偏倚的相对效应最大。简单的无偏方法,并鼓励使用。
It is well known that statistical classification procedures should be assessed using data that are separate from those used to train the classifier. This principle is commonly overlooked when the classification procedure in question is population assignment using a set of genetic markers that were chosen specifically on the basis of their allele frequencies from amongst a larger number of candidate markers. This oversight leads to a systematic upward bias in the predicted accuracy of the chosen set of markers for population assignment. Three widely used software programs for selecting markers informative for population assignment suffer from this bias. The extent of this bias is documented through a small set of simulations. The relative effect of the bias is largest when screening many candidate loci from poorly differentiated populations. Simple unbiased methods are presented and their use encouraged.