FLOCK: a method for quick mapping of admixture without source samples

FLOCK: a method for quick mapping of admixture without source samples
复制标题

DOI:
10.1111/j.1755-0998.2009.02571.x
复制
发表时间:
2009-09-01
影响因子:
7.7
通讯作者:
Turgeon, J.
Turgeon, J.
中科院分区:
生物学1区
文献类型:
--
作者:
Duchesne, P.;Turgeon, J.

文献摘要

被引文献

相似文献

识别和估计个体和/或种群混合是进化和保护生物学中一个非常常见的目标。在许多情况下,来自一个或多个杂交实体的样品不可用或不容易鉴定。在这里,我们描述的FLOCK,一种新的方法,特别是设计用于提供空间和/或时间的混合地图的情况下,一个或几个源样本。FLOCK是一种非贝叶斯方法,因此与以前的聚类算法有很大不同。其工作原理是将所有采集的样本(总样本)重复重新分配到k个子样本中,每次重新分配都比前一次更有效地吸引遗传相似的个体。这种滚雪球效应,更正式地称为正反馈机制,使FLOCK成为一个高效快速的排序过程。FLOCK的使用说明了两个经验的情况下,已彻底分析以前与其他方法。为了更好地评估FLOCK算法的功效,进行了大量的模拟。在FLOCK和Structure算法之间进行了性能比较。当存在不可忽略数量的纯基因型时,两者表现同样好。然而,在没有纯基因型的情况下,FLOCK被证明是更强大的。此外,FLOCK表现出更大的潜力,快速处理。显示运行时间随总样本大小和k大小(进行混合物图谱分析的参比样本数量)线性增加。
Identifying and estimating individual and/or population admixture is a very common objective in evolution and conservation biology. There are many situations where samples from one or many of the putatively hybridizing entities are not available or easily identified. Here we describe FLOCK, a new method especially designed to provide spatial and/or temporal admixture maps in the absence of one or several source samples. FLOCK is a non-Bayesian method and therefore differs substantially from previous clustering algorithms. Its working principle is repeated re-allocation of all collected specimens (total sample) to the k subsamples, each re-allocation being more effective than the previous one in attracting genetically similar individuals. This snowball effect, more formally referred to as a positive feedback mechanism, makes FLOCK an efficient and quick sorting process. The usage of FLOCK is illustrated with two empirical situations which have been thoroughly analysed previously with other approaches. A number of simulations were run to better assess the power of the FLOCK algorithm. Performance comparisons were made between the FLOCK and Structure algorithms. When non-negligible numbers of pure genotypes were present, the two performed equally well. However, FLOCK proved significantly more powerful in the absence of pure genotypes. Moreover, FLOCK showed more potential for fast processing. Run times were shown to increase linearly with size of total sample and with size of k, the number of reference samples from which admixture mapping is performed.