A mixed factors model for dimension reduction and extraction of a group structure in gene expression data

A mixed factors model for dimension reduction and extraction of a group structure in gene expression data
复制标题

DOI:
10.1109/csb.2004.13
复制
发表时间:
2004-08
期刊:
Proceedings. 2004 IEEE Computational Systems Bioinformatics Conference, 2004. CSB 2004.
影响因子:
--
通讯作者:
Ryo Yoshida;T. Higuchi;S. Imoto
Ryo Yoshida;T. Higuchi;S. Imoto
中科院分区:
其他
文献类型:
--
作者:
Ryo Yoshida;T. Higuchi;S. Imoto

文献摘要

被引文献

相似文献

当我们根据基因对组织样本进行聚类时,要分组的观测值数量远小于特征向量的维度。在这种情况下,传统基于模型的聚类的适用性受到限制,因为特征向量的高维数导致密度估计过程中的过度填充。为了克服这一困难,我们尝试对因子分析进行方法论扩展。我们的方法不仅使我们能够防止过度填充的发生,而且还能够处理聚类、数据压缩和提取一组与解释群体结构相关的基因的问题。通过对白血病数据集的应用证明了潜在的有用性。
When we cluster tissue samples on the basis of genes, the number of observations to be grouped is much smaller than the dimension of feature vector. In such a case, the applicability of conventional model-based clustering is limited since the high dimensionality of feature vector leads to overfilling during the density estimation process. To overcome such difficulty, we attempt a methodological extension of the factor analysis. Our approach enables us not only to prevent from the occurrence of overfilling, but also to handle the issues of clustering, data compression and extracting a set of genes to be relevant to explain the group structure. The potential usefulness are demonstrated with the application to the leukemia dataset.