Relative performance of Bayesian clustering software for inferring population substructure and individual assignment at low levels of population differentiation

Relative performance of Bayesian clustering software for inferring population substructure and individual assignment at low levels of population differentiation
复制标题

DOI:
10.1007/s10592-005-9098-1
复制
发表时间:
2006-04-01
影响因子:
2.2
通讯作者:
Rhodes, OE
Rhodes, OE
中科院分区:
环境科学与生态学3区
文献类型:
--
作者:
Latch, EK;Dharmarajan, G;Rhodes, OE

文献摘要

被引文献

相似文献

描述种群间遗传分化的传统方法依赖于个体的先验分组。贝叶斯聚类方法通过使用连锁和Hardy-Weinberg不平衡将个体样本分解为遗传上不同的群体来避免这种限制。有几个软件程序可用于贝叶斯聚类分析,所有这些都描述了随着种群间遗传分化水平的降低,检测不同聚类的能力下降。然而,还没有研究比较这种方法在低水平的种群分化,这可能是常见的物种,其中人口经历了最近的分离或高水平的基因流的性能。我们使用模拟数据来评估三个贝叶斯聚类软件程序,分区,结构和BAPS,在低于F-ST=0.1的人口分化水平的性能。PARTITION无法正确识别亚群的数量,直到F-ST水平达到约0.09。STRUCTURE和BAPS在低水平的种群分化下表现非常好,并且能够在F-ST约为0.03时正确识别亚群的数量。个体基因组分配给其真实来源群体的平均比例随着两个程序的F-ST的增加而增加,在F-ST为0.05时达到92%以上。随着F-ST的增加,错误分配(分配给不正确的亚群)的平均数量继续减少,当F-ST为0.05时,使用任何一种程序,只有不到3%的个体被错误分配。当聚类没有很好地区分时(F-ST=0.02-0.03),STRUCTURE和BAPS都能很好地推断聚类的数量,但我们的结果表明,F-ST必须至少为0.05才能达到大于97%的分配准确度。
Traditional methods for characterizing genetic differentiation among populations rely on a priori grouping of individuals. Bayesian clustering methods avoid this limitation by using linkage and Hardy-Weinberg disequilibrium to decompose a sample of individuals into genetically distinct groups. There are several software programs available for Bayesian clustering analyses, all of which describe a decrease in the ability to detect distinct clusters as levels of genetic differentiation among populations decrease. However, no study has yet compared the performance of such methods at low levels of population differentiation, which may be common in species where populations have experienced recent separation or high levels of gene flow. We used simulated data to evaluate the performance of three Bayesian clustering software programs, PARTITION, STRUCTURE, and BAPS, at levels of population differentiation below F-ST=0.1. PARTITION was unable to correctly identify the number of subpopulations until levels of F-ST reached around 0.09. Both STRUCTURE and BAPS performed very well at low levels of population differentiation, and were able to correctly identify the number of subpopulations at F-ST around 0.03. The average proportion of an individual's genome assigned to its true population of origin increased with increasing F-ST for both programs, reaching over 92% at an F-ST of 0.05. The average number of misassignments (assignments to the incorrect subpopulation) continued to decrease as F-ST increased, and when F-ST was 0.05, fewer than 3% of individuals were misassigned using either program. Both STRUCTURE and BAPS worked extremely well for inferring the number of clusters when clusters were not well-differentiated (F-ST=0.02-0.03), but our results suggest that F-ST must be at least 0.05 to reach an assignment accuracy of greater than 97%.