Clusterflock: a flocking algorithm for isolating congruent phylogenomic datasets.

Clusterflock: a flocking algorithm for isolating congruent phylogenomic datasets.
复制标题

簇氟:用于隔离系统基因组数据集的羊群算法。

DOI:
10.1186/s13742-016-0152-3
复制
发表时间:
2016-10-24
期刊:
影响因子:
9.2
通讯作者:
Planet PJ
Planet PJ
中科院分区:
生物学2区
文献类型:
--
作者:
Narechania A;Baker R;DeSalle R;Mathema B;Kolokotronis SO;Kreiswirth B;Planet PJ

文献摘要

参考文献

被引文献

相似文献

动物的集体行为,如鸟的成群或鱼的浅滩,启发了一类算法,旨在优化各种应用中基于距离的集群,包括文档分析和DNA微阵列。在集群模型中,单个代理只对他们直接的环境做出反应,并根据几个简单的规则移动。在几次迭代之后,代理自组织,簇出现而不需要划分种子。除了无人监督的性质外,群集还提供了几个计算优势,包括可能减少所需的比较次数。在这里提供的工具Clusterflock中,我们实现了一个集群算法,该算法旨在定位具有共同进化历史的直系基因家族(OGFs)的组(群)。成对距离,衡量OGFs引导群形成之间的系统发育不一致性。我们在几个模拟数据集上通过改变底层拓扑的数量、缺失数据的比例和进化速率来测试该方法,结果表明,在包含高度缺失数据和速率异质性的数据集中,Clusterflock的性能优于其他成熟的聚类技术。我们还在金黄色葡萄球菌的一个已知的大规模重组事件中验证了它的实用性。通过分离具有不同系统发育信号的OGFs组,我们能够精确定位重组区域,而无需强制进行预定数量的分组或定义预定的不一致阈值。Clusterflock是一个开源工具,可以用来发现水平转移的基因、染色体的重组区域和基因组的系统发育“核心”。虽然我们在这里使用它是在进化的上下文中,但它可以推广到任何集群问题。用户可以编写扩展程序来计算单位间隔上的任何距离度量,并可以使用这些距离来‘聚集’任何类型的数据。
Collective animal behavior, such as the flocking of birds or the shoaling of fish, has inspired a class of algorithms designed to optimize distance-based clusters in various applications, including document analysis and DNA microarrays. In a flocking model, individual agents respond only to their immediate environment and move according to a few simple rules. After several iterations the agents self-organize, and clusters emerge without the need for partitional seeds. In addition to its unsupervised nature, flocking offers several computational advantages, including the potential to reduce the number of required comparisons. In the tool presented here, Clusterflock, we have implemented a flocking algorithm designed to locate groups (flocks) of orthologous gene families (OGFs) that share an evolutionary history. Pairwise distances that measure phylogenetic incongruence between OGFs guide flock formation. We tested this approach on several simulated datasets by varying the number of underlying topologies, the proportion of missing data, and evolutionary rates, and show that in datasets containing high levels of missing data and rate heterogeneity, Clusterflock outperforms other well-established clustering techniques. We also verified its utility on a known, large-scale recombination event in Staphylococcus aureus. By isolating sets of OGFs with divergent phylogenetic signals, we were able to pinpoint the recombined region without forcing a pre-determined number of groupings or defining a pre-determined incongruence threshold. Clusterflock is an open-source tool that can be used to discover horizontally transferred genes, recombined areas of chromosomes, and the phylogenetic ‘core’ of a genome. Although we used it here in an evolutionary context, it is generalizable to any clustering problem. Users can write extensions to calculate any distance metric on the unit interval, and can use these distances to ‘flock’ any type of data.
DOI: 10.1088/0305-4470/30/5/009
发表时间: 1997-03-07
期刊: JOURNAL OF PHYSICS A-MATHEMATICAL AND GENERAL
影响因子: --
作者:
Czirok, A;Stanley, HE;Vicsek, T
通讯作者: Vicsek, T
DOI: 10.1006/jmbi.1994.1104
发表时间: 1994-02-04
影响因子: 5.6
作者:
KROGH, A;BROWN, M;HAUSSLER, D
通讯作者: HAUSSLER, D
DOI: 10.1006/jtbi.2002.3065
发表时间: 2002-09-07
影响因子: 2
作者:
Couzin, ID;Krause, J;Franks, NR
通讯作者: Franks, NR
DOI: 10.1093/molbev/msr110
发表时间: 2011-10-01
影响因子: 10.7
作者:
Leigh, Jessica W.;Schliep, Klaus;Bapteste, Eric
通讯作者: Bapteste, Eric
DOI: 10.1098/rspb.2013.2450
发表时间: 2014-02-22
影响因子: 4.7
作者:
Boto, Luis
通讯作者: Boto, Luis