Visualization methods for statistical analysis of microarray clusters.

Visualization methods for statistical analysis of microarray clusters.
复制标题

DOI:
10.1186/1471-2105-6-115
复制
发表时间:
2005-05-12
期刊:
影响因子:
3
通讯作者:
Troyanskaya OG
Troyanskaya OG
中科院分区:
生物学4区
文献类型:
--
作者:
Hibbs MA;Dirksen NC;Li K;Troyanskaya OG

文献摘要

参考文献

被引文献

相似文献

在微阵列数据中识别功能相关基因组的最常用方法是应用聚类算法。然而,不可能确定哪种聚类算法最适合应用,并且由于缺乏黄金标准,很难验证任何算法的结果。适当的数据可视化工具可以帮助这个分析过程,但现有的可视化方法没有专门解决这个问题。我们提出了几种可视化技术,将有意义的统计,是噪声鲁棒性的目的,分析结果的聚类算法的微阵列数据。这包括一个基于排名的可视化方法,是更强大的噪音,差异显示方法,以帮助评估集群质量和检测离群值,并将高维数据投影到三维空间,以检查集群之间的关系。我们的方法是交互式的,并动态地连接在一起,以进行全面的分析。此外,我们的方法适用于蛋白质和基因表达微阵列,我们的架构是可扩展的桌面/笔记本电脑屏幕和大规模显示设备上使用。该方法在GeneVAnD(数据集的基因组可视化分析)中实现,可在。将相关的统计信息转化为数据可视化是分析大型生物数据集的关键,特别是因为高水平的噪音和缺乏比较的金标准。我们开发了几种新的可视化技术,并证明了它们的有效性,评估集群质量和集群之间的关系。
The most common method of identifying groups of functionally related genes in microarray data is to apply a clustering algorithm. However, it is impossible to determine which clustering algorithm is most appropriate to apply, and it is difficult to verify the results of any algorithm due to the lack of a gold-standard. Appropriate data visualization tools can aid this analysis process, but existing visualization methods do not specifically address this issue. We present several visualization techniques that incorporate meaningful statistics that are noise-robust for the purpose of analyzing the results of clustering algorithms on microarray data. This includes a rank-based visualization method that is more robust to noise, a difference display method to aid assessments of cluster quality and detection of outliers, and a projection of high dimensional data into a three dimensional space in order to examine relationships between clusters. Our methods are interactive and are dynamically linked together for comprehensive analysis. Further, our approach applies to both protein and gene expression microarrays, and our architecture is scalable for use on both desktop/laptop screens and large-scale display devices. This methodology is implemented in GeneVAnD (Genomic Visual ANalysis of Datasets) and is available at . Incorporating relevant statistical information into data visualizations is key for analysis of large biological datasets, particularly because of high levels of noise and the lack of a gold-standard for comparisons. We developed several new visualization techniques and demonstrated their effectiveness for evaluating cluster quality and relationships between clusters.
DOI: 10.1073/pnas.97.18.10101
发表时间: 2000-08-29
影响因子: 11.1
作者:
Alter, O;Brown, PO;Botstein, D
通讯作者: Botstein, D
DOI: 10.1091/mbc.9.12.3273
发表时间: 1998-12-01
影响因子: 3.3
作者:
Spellman, PT;Sherlock, G;Futcher, B
通讯作者: Futcher, B
DOI: 10.1016/s0014-5793(02)02873-9
发表时间: 2002-07-03
期刊: FEBS LETTERS
影响因子: 3.5
作者:
Méndez, MA;Hödar, C;Cambiazo, V
通讯作者: Cambiazo, V
DOI: 10.1093/bioinformatics/btg136
发表时间: 2003-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Johnson, JE;Stromvik, MV;Retzel, EF
通讯作者: Retzel, EF
DOI: 10.2144/03342mt01
发表时间: 2003-02-01
期刊: BIOTECHNIQUES
影响因子: 2.7
作者:
Saeed, AI;Sharov, V;Quackenbush, J
通讯作者: Quackenbush, J