Wavelet-Based Clustering for Mixed-Effects Functional Models in High Dimension

Wavelet-Based Clustering for Mixed-Effects Functional Models in High Dimension
复制标题

DOI:
10.1111/j.1541-0420.2012.01828.x
复制
发表时间:
2013-03-01
期刊:
影响因子:
1.9
通讯作者:
Picard, F.
Picard, F.
中科院分区:
数学3区
文献类型:
--
作者:
Giacofci, M.;Lambert-Lacroix, S.;Picard, F.

文献摘要

被引文献

相似文献

我们提出了一种方法,高维曲线聚类存在个体间变异。长期以来,人们一直在研究曲线聚类,特别是使用样条来解释函数随机效应。然而,样条函数在处理高维数据时并不合适,并且不能用于建模不规则曲线,如峰状数据。我们的方法是基于固定和随机效应的信号的小波分解。我们提出了一个有效的降维步骤的基础上小波阈值适应于多个曲线,并使用一个适当的结构的随机效应方差,我们确保固定和随机效应位于同一个功能空间,即使在处理不规则的功能,属于Besov空间。在小波域中,我们的模型恢复到线性混合效应模型,可以用于基于模型的聚类算法,我们开发了EM算法的最大似然估计。通过广泛的模拟研究验证了整个过程的属性。然后,我们说明了我们的方法质谱数据,我们提出了一个原始的应用程序的功能数据分析微阵列比较基因组杂交(CGH)数据。我们的程序可以通过R包curvature获得,这是第一个公开可用的包,它在高维框架(可在CRAN上获得)中执行具有随机效应的曲线聚类。
We propose a method for high-dimensional curve clustering in the presence of interindividual variability. Curve clustering has longly been studied especially using splines to account for functional random effects. However, splines are not appropriate when dealing with high-dimensional data and can not be used to model irregular curves such as peak-like data. Our method is based on a wavelet decomposition of the signal for both fixed and random effects. We propose an efficient dimension reduction step based on wavelet thresholding adapted to multiple curves and using an appropriate structure for the random effect variance, we ensure that both fixed and random effects lie in the same functional space even when dealing with irregular functions that belong to Besov spaces. In the wavelet domain our model resumes to a linear mixed-effects model that can be used for a model-based clustering algorithm and for which we develop an EM-algorithm for maximum likelihood estimation. The properties of the overall procedure are validated by an extensive simulation study. Then, we illustrate our method on mass spectrometry data and we propose an original application of functional data analysis on microarray comparative genomic hybridization (CGH) data. Our procedure is available through the R package curvclust which is the first publicly available package that performs curve clustering with random effects in the high dimensional framework (available on the CRAN).