A systematic comparison of genome-scale clustering algorithms.

A systematic comparison of genome-scale clustering algorithms.
复制标题

DOI:
10.1186/1471-2105-13-s10-s7
复制
发表时间:
2012-06-25
期刊:
影响因子:
3
通讯作者:
Langston MA
Langston MA
中科院分区:
生物学4区
文献类型:
--
作者:
Jay JJ;Eblen JD;Zhang Y;Benson M;Perkins AD;Saxton AM;Voy BH;Chesler EJ;Langston MA

文献摘要

被引文献

相似文献

大量的聚类算法已经被应用到基因共表达实验中。这些算法涵盖了广泛的方法,从传统的技术,如k-means和层次聚类,图形化的方法,如k-团社区,加权基因共表达网络(WGCNA)和paraclique。比较这些方法的相对有效性,为算法的选择、开发和实施提供指导。大多数以前的比较聚类评价工作集中在参数方法。图论方法是最近添加到工具集的全球分析和分解的微阵列共表达矩阵,一般不包括在早期的方法比较。在本研究中,各种参数和图论聚类算法进行了比较,使用良好的特征转录组数据在基因组规模从酿酒酵母。对于研究中的每种聚类方法,测试了各种参数。使用Jaccard相似性来测量每个聚类与每个GO和KEGG注释集的一致性,并且将最高Jaccard得分分配给聚类。将聚类分组为小、中、大箱,并将每个箱中前五个评分聚类的Jaccard评分平均,并报告为特定方法的最佳平均前5名(BAT 5)评分。基于与已知途径的阳性匹配,对每种方法产生的簇进行评价。这就产生了一个易于解释的基因聚类相对有效性的排名。方法也进行了测试,以确定他们是否能够识别集群与其他聚类方法确定的一致。对已知基因分类的聚类验证表明,对于这些数据,基于图形的技术优于传统的聚类方法,这表明进一步开发和应用组合策略是必要的。
A wealth of clustering algorithms has been applied to gene co-expression experiments. These algorithms cover a broad range of approaches, from conventional techniques such as k-means and hierarchical clustering, to graphical approaches such as k-clique communities, weighted gene co-expression networks (WGCNA) and paraclique. Comparison of these methods to evaluate their relative effectiveness provides guidance to algorithm selection, development and implementation. Most prior work on comparative clustering evaluation has focused on parametric methods. Graph theoretical methods are recent additions to the tool set for the global analysis and decomposition of microarray co-expression matrices that have not generally been included in earlier methodological comparisons. In the present study, a variety of parametric and graph theoretical clustering algorithms are compared using well-characterized transcriptomic data at a genome scale from Saccharomyces cerevisiae. For each clustering method under study, a variety of parameters were tested. Jaccard similarity was used to measure each cluster's agreement with every GO and KEGG annotation set, and the highest Jaccard score was assigned to the cluster. Clusters were grouped into small, medium, and large bins, and the Jaccard score of the top five scoring clusters in each bin were averaged and reported as the best average top 5 (BAT5) score for the particular method. Clusters produced by each method were evaluated based upon the positive match to known pathways. This produces a readily interpretable ranking of the relative effectiveness of clustering on the genes. Methods were also tested to determine whether they were able to identify clusters consistent with those identified by other clustering methods. Validation of clusters against known gene classifications demonstrate that for this data, graph-based techniques outperform conventional clustering approaches, suggesting that further development and application of combinatorial strategies is warranted.