Judging the quality of gene expression-based clustering methods using gene annotation

Judging the quality of gene expression-based clustering methods using gene annotation
复制标题

DOI:
10.1101/gr.397002
复制
发表时间:
2002-10-01
期刊:
影响因子:
7
通讯作者:
Roth, FP
Roth, FP
中科院分区:
生物学1区
文献类型:
--
作者:
Gibbons, FD;Roth, FP

文献摘要

被引文献

相似文献

我们比较了几种常用的表达为基础的基因聚类算法使用的品质因数的基础上的聚类成员和已知的基因属性之间的互信息。通过研究各种公开的表达数据集,我们得出结论,富集的生物功能的集群,一般来说,在相当低的集群数最高。作为两个基因的表达模式之间的相异性的测量,没有方法优于基于比率的测量的欧几里得距离,或在聚类数的最佳选择的非基于比率的测量的皮尔逊距离。我们表明,自组织地图的方法是最好的两种测量类型在更高数量的集群。来自单连锁和平均连锁层次聚类的基因簇往往产生比随机结果更差的结果。
We compare several commonly used expression-based gene clustering algorithms using a figure of merit based on the mutual information between cluster membership and known gene attributes. By studying various publicly available expression data sets we conclude that enrichment of clusters for biological function is, in general, highest at rather low cluster numbers. As a measure of dissimilarity between the expression patterns of two genes, no method outperforms Euclidean distance for ratio-based measurements, or Pearson distance for non-ratio-based measurements at the optimal choice of cluster number. We show the self-organized-map approach to be best for both measurement types at higher numbers of clusters. Clusters of genes derived from single- and average-linkage hierarchical clustering tend to produce worse-than-random results.