A weighted Mutual Information Biclustering algorithm for gene expression data

A weighted Mutual Information Biclustering algorithm for gene expression data
复制标题

DOI:
10.2298/csis170301021y
复制
发表时间:
2017
期刊:
Comput. Sci. Inf. Syst.
影响因子:
--
通讯作者:
Yidong Li;Wenhua Liu;Y. Jia;Hai-rong Dong
Yidong Li;Wenhua Liu;Y. Jia;Hai-rong Dong
中科院分区:
其他
文献类型:
--
作者:
Yidong Li;Wenhua Liu;Y. Jia;Hai-rong Dong

文献摘要

相似文献

基因芯片是实验分子生物学的最新突破之一,已经提供了大量的高维遗传数据。传统的聚类方法很难处理这种高维数据,其一个子集的基因在一个子集的条件下的共同调控。双聚类算法被引入到发现基因表达数据的局部特征。本文提出了一种新的双聚类算法--加权互信息双聚类算法(WMIB),用于发现基因表达数据的这种局部特征。该算法采用加权互信息作为新的相似性度量,可以同时检测基因间复杂的线性和非线性关系,并提出了一种新的目标函数来更新每个双聚类的权值,该目标函数可以根据一定的规则同时选择每个双聚类的条件集。实验结果表明,该算法可以同时生成较大的双聚类和较小的均方残差。
Microarrays are one of the latest breakthroughs in experimental molecular biology, which have already provided huge amount of high dimensional genetic data. Traditional clustering methods are difficult to deal with this high dimensional data, whose a subset of genes are co-regulated under a subset of conditions. Biclustering algorithms are introduced to discover local characteristics of gene expression data. In this paper, we present a novel biclustering algorithm, which calledWeighted Mutual Information Biclustering algorithm (WMIB) to discover this local characteristics of gene expression data. In our algorithm, we use the weighted mutual information as new similarity measure which can be simultaneously detect complex linear and nonlinear relationships between genes, and our algorithm proposes a new objective function to update weights of each bicluster, which can simultaneously select the conditions set of each bicluster using some rules.We have evaluated our algorithm on yeast gene expression data, the experimental results show that our algorithm can generate larger biclusters with lower mean square residues simultaneously.