Mixture of linear mixed models for clustering gene expression profiles from repeated microarray experiments

Mixture of linear mixed models for clustering gene expression profiles from repeated microarray experiments
复制标题

DOI:
10.1191/1471082x05st096oa
复制
发表时间:
2005-10-01
影响因子:
1
通讯作者:
Lavergne, C
Lavergne, C
中科院分区:
数学4区
文献类型:
--
作者:
Celeux, G;Martin, O;Lavergne, C

文献摘要

被引文献

相似文献

数据变异性在微阵列数据分析中可能是重要的。因此,当聚类基因表达谱时,利用重复数据可能是明智的。本文研究了基于模型的聚类分析中重复数据的分析问题。选择线性混合模型以考虑数据变异性,并考虑这些模型的混合。这导致了一个大范围的可能的模型,这取决于对观测值的协方差结构和混合模型的假设。给出了这类模型的EM算法极大似然估计。选择一个特定的混合线性混合模型的问题被认为是使用惩罚似然准则。说明性的蒙特卡罗实验和应用程序的基因表达谱的聚类进行了详细说明。所有这些实验都突出了线性混合模型混合物的兴趣,以考虑聚类分析上下文中的数据变异性。
Data variability can be important in microarray data analysis. Thus, when clustering gene expression profiles, it could be judicious to make use of repeated data. In this paper, the problem of analysing repeated data in the model-based cluster analysis context is considered. Linear mixed models are chosen to take into account data variability and mixture of these models are considered. This leads to a large range of possible models depending on the assumptions made on both the covariance structure of the observations and the mixture model. The maximum likelihood estimation of this family of models through the EM algorithm is presented. The problem of selecting a particular mixture of linear mixed models is considered using penalized likelihood criteria. Illustrative Monte Carlo experiments are presented and an application to the clustering of gene expression profiles is detailed. All those experiments highlight the interest of linear mixed model mixtures to take into account data variability in a cluster analysis context.