TreeCluster: Clustering biological sequences using phylogenetic trees

TreeCluster: Clustering biological sequences using phylogenetic trees
复制标题

DOI:
10.1371/journal.pone.0221068
复制
发表时间:
2019-08-22
期刊:
影响因子:
3.7
通讯作者:
Mirarab, Siavash
Mirarab, Siavash
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Balaban, Metin;Moshiri, Niema;Mirarab, Siavash

文献摘要

被引文献

相似文献

基于同源序列的相似性对同源序列进行聚类是许多生物信息学应用中出现的问题。序列聚类的事实最终是它们的系统发育关系的结果。尽管有这种观察和树可以定义聚类的自然方式,但序列聚类的大多数应用并不使用系统发育树,而是对成对序列距离进行操作。由于大规模系统发育推断的进展,我们认为,基于树的聚类是利用不足。我们定义了一个家庭的优化问题,给定一个任意的树,返回最小数量的集群,使所有集群坚持其异质性的约束。我们研究了三个特定的约束,限制(1)每个簇的直径,(2)其分支长度的总和,或(3)成对距离的链。这三个问题可以在随树的大小线性增加的时间内解决,并且对于三个标准中的两个,算法在理论计算机科学家文献中是已知的。我们在一个名为TreeCluster的工具中实现了这些算法,我们在三个应用程序上进行了测试:微生物组数据的OTU聚类,HIV传播聚类和分治多序列比对。我们表明,通过使用基于树的距离,TreeCluster生成更多的内部一致的集群比替代品,并提高下游应用程序的有效性。TreeCluster可以在https://github.cominiernascifireetcher上找到。
Clustering homologous sequences based on their similarity is a problem that appears in many bioinformatics applications. The fact that sequences cluster is ultimately the result of their phylogenetic relationships. Despite this observation and the natural ways in which a tree can define clusters, most applications of sequence clustering do not use a phylogenetic tree and instead operate on pairwise sequence distances. Due to advances in large-scale phylogenetic inference, we argue that tree-based clustering is under-utilized. We define a family of optimization problems that, given an arbitrary tree, return the minimum number of clusters such that all clusters adhere to constraints on their heterogeneity. We study three specific constraints, limiting (1) the diameter of each cluster, (2) the sum of its branch lengths, or (3) chains of pairwise distances. These three problems can be solved in time that increases linearly with the size of the tree, and for two of the three criteria, the algorithms have been known in the theoretical computer scientist literature. We implement these algorithms in a tool called TreeCluster, which we test on three applications: OTU clustering for microbiome data, HIV transmission clustering, and divide-and-conquer multiple sequence alignment. We show that, by using tree-based distances, TreeCluster generates more internally consistent clusters than alternatives and improves the effectiveness of downstream applications. TreeCluster is available at https://github.cominiernascifireeCluster.