Scoring clustering solutions by their biological relevance

Scoring clustering solutions by their biological relevance
复制标题

DOI:
10.1093/bioinformatics/btg330
复制
发表时间:
2003-12-12
期刊:
影响因子:
5.8
通讯作者:
Shamir, R
Shamir, R
中科院分区:
生物学3区
文献类型:
--
作者:
Gat-Viks, I;Sharan, R;Shamir, R

文献摘要

被引文献

相似文献

动机:基因表达数据分析的中心步骤是识别表现出相似表达模式的基因组。将基因表达数据分成同质组在功能注释、组织分类、调控基序识别和其他应用中被证明是有用的。虽然有丰富的文献关于基因表达分析的聚类算法,但很少有工作涉及到对聚类结果的系统比较和评估。通常,不同的聚类算法会在相同的数据上产生不同的聚类解,并且没有一致的准则来选择它们。结果:我们提出了一种新的基于统计学的方法来根据先验的生物学知识来评估聚类解。我们的方法可以用来比较不同的聚类方案或优化聚类算法的参数。该方法基于将聚类元素的生物属性向量投影到实线上,使得组间和组内方差估计器的比率最大化。然后,使用非参数方差分析检验对预测数据进行评分,并评估评分的置信度。我们使用模拟数据对我们的方法进行了验证,结果表明我们的评分方法优于现有的几种方法,包括分离度与同质性比和轮廓度量。我们应用我们的方法对酵母细胞周期基因表达数据的几种聚类方法的结果进行了评估。
Motivation: A central step in the analysis of gene expression data is the identification of groups of genes that exhibit similar expression patterns. Clustering gene expression data into homogeneous groups was shown to be instrumental in functional annotation, tissue classification, regulatory motif identification, and other applications. Although there is a rich literature on clustering algorithms for gene expression analysis, very few works addressed the systematic comparison and evaluation of clustering results. Typically, different clustering algorithms yield different clustering solutions on the same data, and there is no agreed upon guideline for choosing among them.Results: We developed a novel statistically based method for assessing a clustering solution according to prior biological knowledge. Our method can be used to compare different clustering solutions or to optimize the parameters of a clustering algorithm. The method is based on projecting vectors of biological attributes of the clustered elements onto the real line, such that the ratio of between-groups and within-group variance estimators is maximized. The projected data are then scored using a non-parametric analysis of variance test, and the score's confidence is evaluated. We validate our approach using simulated data and show that our scoring method outperforms several extant methods, including the separation to homogeneity ratio and the silhouette measure. We apply our method to evaluate results of several clustering methods on yeast cell-cycle gene expression data.