Biclustering algorithms for biological data analysis: A survey

Biclustering algorithms for biological data analysis: A survey
复制标题

DOI:
10.1109/tcbb.2004.2
复制
发表时间:
2004-01-01
影响因子:
4.5
通讯作者:
Oliveira, AL
Oliveira, AL
中科院分区:
工程技术3区
文献类型:
--
作者:
Madeira, SC;Oliveira, AL

文献摘要

被引文献

相似文献

已经提出了大量的聚类方法来分析从微阵列实验获得的基因表达数据。但是,标准聚类方法应用于基因的结果受到限制。这种限制是由于存在基因活性不相关的许多实验条件所施加的。进行条件聚类时,存在类似的限制。因此,已经提出了许多在数据矩阵的行和列尺寸上同时聚类的算法。目的是找到一量,即基因的亚组和条件的亚组,其中基因在每种情况下都表现出高度相关的活性。在本文中,我们将此类别的算法称为双簇。在文献中也将双簇称为共簇和直接聚类,以及其他名称,并且也已用于信息检索和数据挖掘等领域。在这项综合调查中,我们分析了许多现有的双簇方法,并根据他们可以找到的双晶浮器类型对它们进行分类,发现了发现的双群体模式,用于执行搜索的方法,用于执行搜索的方法,用于执行的方法,用于执行搜索的方法评估解决方案和目标应用程序。
A large number of clustering approaches have been proposed for the analysis of gene expression data obtained from microarray experiments. However, the results from the application of standard clustering methods to genes are limited. This limitation is imposed by the existence of a number of experimental conditions where the activity of genes is uncorrelated. A similar limitation exists when clustering of conditions is performed. For this reason, a number of algorithms that perform simultaneous clustering on the row and column dimensions of the data matrix has been proposed. The goal is to find submatrices, that is, subgroups of genes and subgroups of conditions, where the genes exhibit highly correlated activities for every condition. In this paper, we refer to this class of algorithms as biclustering. Biclustering is also referred in the literature as coclustering and direct clustering, among others names, and has also been used in fields such as information retrieval and data mining. In this comprehensive survey, we analyze a large number of existing approaches to biclustering, and classify them in accordance with the type of biclusters they can find, the patterns of biclusters that are discovered, the methods used to perform the search, the approaches used to evaluate the solution, and the target applications.