Detecting the number of clusters of individuals using the software STRUCTURE: a simulation study

Detecting the number of clusters of individuals using the software STRUCTURE: a simulation study
复制标题

DOI:
10.1111/j.1365-294x.2005.02553.x
复制
发表时间:
2005-07-01
期刊:
影响因子:
4.9
通讯作者:
Goudet, J
Goudet, J
中科院分区:
生物学1区
文献类型:
--
作者:
Evanno, G;Regnaut, S;Goudet, J

文献摘要

被引文献

相似文献

遗传同质群体的识别是群体遗传学中一个长期存在的问题。最近在软件结构中实现的贝叶斯算法允许识别这样的群体。然而,这种算法的能力,以检测真正的集群数(K)在一个样本中的个人时,人口之间的扩散模式是不均匀的还没有被测试。本研究的目的是进行这样的测试,使用各种分散的情况下,从数据生成的基于个人的模型。我们发现,在大多数情况下,估计的“数据的对数概率”并不能提供聚类数K的正确估计。然而,使用基于连续K值之间数据的对数概率变化率的特定统计量Delta K,我们发现STRUCTURE可以准确地检测我们测试的场景的最高层次结构。正如预期的那样,结果对所使用的遗传标记的类型(AFLP与微卫星),位点评分的数量,采样的种群数量以及每个样本中分型的个体数量很敏感。
The identification of genetically homogeneous groups of individuals is a long standing issue in population genetics. A recent Bayesian algorithm implemented in the software STRUCTURE allows the identification of such groups. However, the ability of this algorithm to detect the true number of clusters (K) in a sample of individuals when patterns of dispersal among populations are not homogeneous has not been tested. The goal of this study is to carry out such tests, using various dispersal scenarios from data generated with an individual-based model. We found that in most cases the estimated 'log probability of data' does not provide a correct estimation of the number of clusters, K. However, using an ad hoc statistic Delta K based on the rate of change in the log probability of data between successive K values, we found that STRUCTURE accurately detects the uppermost hierarchical level of structure for the scenarios we tested. As might be expected, the results are sensitive to the type of genetic marker used (AFLP vs. microsatellite), the number of loci scored, the number of populations sampled, and the number of individuals typed in each sample.