The Gap Procedure: for the identification of phylogenetic clusters in HIV-1 sequence data.

The Gap Procedure: for the identification of phylogenetic clusters in HIV-1 sequence data.
复制标题

DOI:
10.1186/s12859-015-0791-x
复制
发表时间:
2015-11-04
期刊:
影响因子:
3
通讯作者:
Brenner BG
Brenner BG
中科院分区:
生物学4区
文献类型:
--
作者:
Vrbik I;Stephens DA;Roger M;Brenner BG

文献摘要

被引文献

相似文献

在传染病的背景下,序列聚类可用于提供对传播动态的重要见解。聚类分析通常使用系统发育方法进行,其中聚类基于足够小的遗传距离和高的自举支持(或后验概率)进行分配。这种系统发育阈值方法的计算负担是一个主要缺点,特别是当考虑大量序列时。此外,该方法需要熟练的用户指定适当的阈值,该阈值可以根据应用而广泛变化。本文提出了差距程序,一个基于距离的聚类算法的DNA序列的分类从感染人类免疫缺陷病毒1型(HIV-1)的个人。我们的启发式算法绕过了系统发育重建的需要,从而支持大型遗传数据集的快速分析。此外,这个完全自动化的过程依赖于排序的成对距离中的数据驱动的间隙来推断聚类,因此不需要用户指定的阈值。通过差距程序对真实的和模拟数据得到的聚类结果与使用阈值方法得到的聚类结果非常一致,而只需要一小部分时间来完成分析。除了在计算时间上的显著增益之外,差距程序在发现遗传相似序列的不同组方面非常有效,并且消除了对主观用户指定值的需要。通过这一程序返回的遗传相似序列簇可用于检测HIV-1传播模式,从而有助于预防、治疗和遏制该疾病。本文的在线版本(doi:10.1186/s12859-015-0791-x)包含补充材料,可供授权用户使用。
In the context of infectious disease, sequence clustering can be used to provide important insights into the dynamics of transmission. Cluster analysis is usually performed using a phylogenetic approach whereby clusters are assigned on the basis of sufficiently small genetic distances and high bootstrap support (or posterior probabilities). The computational burden involved in this phylogenetic threshold approach is a major drawback, especially when a large number of sequences are being considered. In addition, this method requires a skilled user to specify the appropriate threshold values which may vary widely depending on the application. This paper presents the Gap Procedure, a distance-based clustering algorithm for the classification of DNA sequences sampled from individuals infected with the human immunodeficiency virus type 1 (HIV-1). Our heuristic algorithm bypasses the need for phylogenetic reconstruction, thereby supporting the quick analysis of large genetic data sets. Moreover, this fully automated procedure relies on data-driven gaps in sorted pairwise distances to infer clusters, thus no user-specified threshold values are required. The clustering results obtained by the Gap Procedure on both real and simulated data, closely agree with those found using the threshold approach, while only requiring a fraction of the time to complete the analysis. Apart from the dramatic gains in computational time, the Gap Procedure is highly effective in finding distinct groups of genetically similar sequences and obviates the need for subjective user-specified values. The clusters of genetically similar sequences returned by this procedure can be used to detect patterns in HIV-1 transmission and thereby aid in the prevention, treatment and containment of the disease. The online version of this article (doi:10.1186/s12859-015-0791-x) contains supplementary material, which is available to authorized users.