SC3: Triple Spectral Clustering- Based Consensus Clustering Framework for Class Discovery from Cancer Gene Expression Profiles

SC3: Triple Spectral Clustering- Based Consensus Clustering Framework for Class Discovery from Cancer Gene Expression Profiles
复制标题

SC3:基于三重谱聚类的共识聚类框架,用于从癌症基因表达谱中发现类别

DOI:
10.1109/tcbb.2012.108
复制
发表时间:
2012-11-01
影响因子:
4.5
通讯作者:
Han, Guoqiang
Han, Guoqiang
中科院分区:
工程技术3区
文献类型:
--
作者:
Yu, Zhiwen;Li, Le;Han, Guoqiang

文献摘要

被引文献

相似文献

为了成功地诊断和治疗癌症,发现和正确分类癌症类型是必不可少的。从癌症数据集中进行类发现的一个具有挑战性的特性是癌症基因表达谱不仅包括大量的基因,而且还包含大量的噪声基因。为了减少噪声基因对癌症基因表达谱的影响,本文提出了两种新的一致性聚类框架:基于三谱聚类的一致性聚类(SC 3)和基于双谱聚类的一致性聚类(SC(2)Ncut),用于从基因表达谱中发现癌症。SC3将谱聚类(SC)算法多次集成到集成框架中以处理基因表达谱。具体地,谱聚类被应用于对基因维度和癌症样本维度进行聚类,并且还被用作共识函数来划分由多个聚类解构造的共识矩阵。与SC 3相比,SC(2)Ncut采用归一化割算法代替谱聚类作为一致性函数。在合成数据集和真实的癌症基因表达谱上的实验表明,该方法不仅在基因表达谱上取得了良好的性能,而且在从这些基因表达谱中发现类别的过程中,性能优于大多数现有方法.
In order to perform successful diagnosis and treatment of cancer, discovering, and classifying cancer types correctly is essential. One of the challenging properties of class discovery from cancer data sets is that cancer gene expression profiles not only include a large number of genes, but also contains a lot of noisy genes. In order to reduce the effect of noisy genes in cancer gene expression profiles, we propose two new consensus clustering frameworks, named as triple spectral clustering-based consensus clustering (SC3) and double spectral clustering-based consensus clustering (SC(2)Ncut) in this paper, for cancer discovery from gene expression profiles. SC3 integrates the spectral clustering (SC) algorithm multiple times into the ensemble framework to process gene expression profiles. Specifically, spectral clustering is applied to perform clustering on the gene dimension and the cancer sample dimension, and also used as the consensus function to partition the consensus matrix constructed from multiple clustering solutions. Compared with SC3, SC(2)Ncut adopts the normalized cut algorithm, instead of spectral clustering, as the consensus function. Experiments on both synthetic data sets and real cancer gene expression profiles illustrate that the proposed approaches not only achieve good performance on gene expression profiles, but also outperforms most of the existing approaches in the process of class discovery from these profiles.