Coral: an integrated suite of visualizations for comparing clusterings.

Coral: an integrated suite of visualizations for comparing clusterings.
复制标题

DOI:
10.1186/1471-2105-13-276
复制
发表时间:
2012-10-29
期刊:
影响因子:
3
通讯作者:
Kingsford C
Kingsford C
中科院分区:
生物学4区
文献类型:
--
作者:
Filippova D;Gadani A;Kingsford C

文献摘要

参考文献

被引文献

相似文献

聚类已经成为许多类型的生物数据(例如相互作用网络,基因表达,宏基因组丰度)的标准分析。在实践中,通过改变使用哪种聚类算法、考虑哪些数据属性、如何设置算法参数以及选择哪些接近最优的聚类,可以获得大量矛盾的聚类。这是一个困难的任务,筛选,虽然这样一个大的集合不同的聚类,以确定哪些聚类功能的参数设置的影响,或工件的特定算法,并表示有意义的模式。了解哪些项经常聚集在一起有助于提高我们对底层数据的理解,并增加我们对生成模块的信心。我们提出珊瑚,一个应用程序的集群大合奏互动式探索。Coral使所有对所有的聚类比较变得容易,支持探索单个聚类,允许跨聚类跟踪模块,并支持识别模块中的核心和外围项目。我们将讨论Coral中的每个可视化组件如何处理与聚类比较相关的特定问题,并提供它们的使用示例。我们还展示了如何珊瑚可以用来直观和定量地比较聚类与地面真相聚类。作为一个案例研究,我们比较最近发表的蛋白质相互作用网络的拟南芥的聚类。我们使用几种流行的算法来生成网络的聚类。我们发现,聚类变化显着,很少有蛋白质是一贯的共同聚集在所有聚类。这证明了在评估基因、蛋白质或序列模块时通常应考虑几个聚类,并且Coral可用于对这些聚类集合进行全面分析。
Clustering has become a standard analysis for many types of biological data (e.g interaction networks, gene expression, metagenomic abundance). In practice, it is possible to obtain a large number of contradictory clusterings by varying which clustering algorithm is used, which data attributes are considered, how algorithmic parameters are set, and which near-optimal clusterings are chosen. It is a difficult task to sift though such a large collection of varied clusterings to determine which clustering features are affected by parameter settings or are artifacts of particular algorithms and which represent meaningful patterns. Knowing which items are often clustered together helps to improve our understanding of the underlying data and to increase our confidence about generated modules. We present Coral, an application for interactive exploration of large ensembles of clusterings. Coral makes all-to-all clustering comparison easy, supports exploration of individual clusterings, allows tracking modules across clusterings, and supports identification of core and peripheral items in modules. We discuss how each visual component in Coral tackles a specific question related to clustering comparison and provide examples of their use. We also show how Coral could be used to visually and quantitatively compare clusterings with a ground truth clustering. As a case study, we compare clusterings of a recently published protein interaction network of Arabidopsis thaliana. We use several popular algorithms to generate the network’s clusterings. We find that the clusterings vary significantly and that few proteins are consistently co-clustered in all clusterings. This is evidence that several clusterings should typically be considered when evaluating modules of genes, proteins, or sequences, and Coral can be used to perform a comprehensive analysis of these clustering ensembles.
DOI: 10.1186/1752-0509-4-100
发表时间: 2010-07-22
影响因子: --
作者:
Lewis AC;Jones NS;Porter MA;Deane CM
通讯作者: Deane CM
DOI: 10.1371/journal.pcbi.1001057
发表时间: 2011-01-20
影响因子: 4.3
作者:
Langfelder P;Luo R;Oldham MC;Horvath S
通讯作者: Horvath S
DOI: 10.1186/1471-2105-6-115
发表时间: 2005-05-12
期刊: BMC bioinformatics
影响因子: 3
作者:
Hibbs MA;Dirksen NC;Li K;Troyanskaya OG
通讯作者: Troyanskaya OG
DOI: 10.1073/pnas.0601602103
发表时间: 2006-06-06
影响因子: 11.1
作者:
Newman, M. E. J.
通讯作者: Newman, M. E. J.
DOI: 10.1093/nar/gkq1157
发表时间: 2011-01
影响因子: 14.9
作者:
Mewes HW;Ruepp A;Theis F;Rattei T;Walter M;Frishman D;Suhre K;Spannagl M;Mayer KF;Stümpflen V;Antonov A
通讯作者: Antonov A