Clustering gene expression data with a penalized graph-based metric.

Clustering gene expression data with a penalized graph-based metric.
复制标题

DOI:
10.1186/1471-2105-12-2
复制
发表时间:
2011-01-04
期刊:
影响因子:
3
通讯作者:
Granitto PM
Granitto PM
中科院分区:
生物学4区
文献类型:
--
作者:
Bayá AE;Granitto PM

文献摘要

参考文献

被引文献

相似文献

在微阵列数据集中寻找簇结构是所谓的“组学”的基本问题。聚类中的一个困难问题是如何处理具有流形结构的数据,即不是以点的紧凑云的形式成形的数据,形成嵌入在高维空间中的任意形状或路径,如可能是某些基因表达数据集的情况。在这项工作中,我们介绍了惩罚k-最近邻图(PKNNG)为基础的度量,在这种情况下,一个新的工具来评估距离。新的度量可以与大多数聚类算法结合使用。PKNNG度量基于两步过程:首先,它使用低k值构建感兴趣数据集的k最近邻图,然后添加具有高度惩罚权重的边,用于连接第一步产生的子图。我们讨论了几种可能的方案连接不同的子图以及惩罚函数。我们展示了几个公共基因表达数据集和模拟人工问题的聚类结果,以评估新度量的行为。在所有情况下,PKNNG度量都显示出有希望的聚类结果。PKNNG度量的使用可以将常用的基于成对距离的聚类方法的性能提高到更高级算法的水平。新方法的一个很大的优点是,研究人员不需要学习新方法,他们可以简单地使用PKNNG度量计算距离,然后,例如,使用层次聚类来生成高维数据的准确和高度可解释的树状图。
The search for cluster structure in microarray datasets is a base problem for the so-called "-omic sciences". A difficult problem in clustering is how to handle data with a manifold structure, i.e. data that is not shaped in the form of compact clouds of points, forming arbitrary shapes or paths embedded in a high-dimensional space, as could be the case of some gene expression datasets. In this work we introduce the Penalized k-Nearest-Neighbor-Graph (PKNNG) based metric, a new tool for evaluating distances in such cases. The new metric can be used in combination with most clustering algorithms. The PKNNG metric is based on a two-step procedure: first it constructs the k-Nearest-Neighbor-Graph of the dataset of interest using a low k-value and then it adds edges with a highly penalized weight for connecting the subgraphs produced by the first step. We discuss several possible schemes for connecting the different sub-graphs as well as penalization functions. We show clustering results on several public gene expression datasets and simulated artificial problems to evaluate the behavior of the new metric. In all cases the PKNNG metric shows promising clustering results. The use of the PKNNG metric can improve the performance of commonly used pairwise-distance based clustering methods, to the level of more advanced algorithms. A great advantage of the new procedure is that researchers do not need to learn a new method, they can simply compute distances with the PKNNG metric and then, for example, use hierarchical clustering to produce an accurate and highly interpretable dendrogram of their high-dimensional data.
在412例患者的基于人群的队列中,乳腺癌的固有分子特征。
DOI: 10.1186/bcr1517
发表时间: 2006
影响因子: 7.4
作者:
Calza, Stefano;Hall, Per;Auer, Gert;Bjohle, Judith;Klaar, Sigrid;Kronenwett, Ulrike;T Liu, Edison;Miller, Lance;Ploner, Alexander;Smeds, Johanna;Bergh, Jonas;Pawitan, Yudi
通讯作者: Pawitan, Yudi
Multi-K:使用集合K均值聚类对微阵列亚型进行准确分类。
DOI: 10.1186/1471-2105-10-260
发表时间: 2009-08-22
期刊: BMC bioinformatics
影响因子: 3
作者:
Kim EY;Kim SY;Ashlock D;Nam D
通讯作者: Nam D
DOI: 10.1109/tpami.2005.113
发表时间: 2005-06-01
影响因子: 23.6
作者:
Fred, ALN;Jain, AK
通讯作者: Jain, AK
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1186/gb-2005-6-10-r88
发表时间: 2005
期刊: Genome biology
影响因子: 12.3
作者:
Dettling M;Gabrielson E;Parmigiani G
通讯作者: Parmigiani G