Functional data analysis for sparse longitudinal data

Functional data analysis for sparse longitudinal data
复制标题

DOI:
10.1198/016214504000001745
复制
发表时间:
2005-06-01
影响因子:
3.7
通讯作者:
Wang, JL
Wang, JL
中科院分区:
数学1区
文献类型:
--
作者:
Yao, F;Müller, HG;Wang, JL

文献摘要

被引文献

相似文献

我们提出了一种非参数的方法来执行稀疏纵向数据的情况下,功能主成分分析。该方法的目的是不规则间隔的纵向数据,其中每个主题的重复测量的数量是小的。相比之下,经典的功能数据分析需要大量的定期间隔测量每个主题。我们假设,重复的测量是随机定位的,每个主题的重复随机数,并确定了一个潜在的平滑随机(特定于主题)轨迹加上测量误差。我们的方法的基本要素是简约估计的协方差结构和平均函数的轨迹,和估计的方差的测量误差。本征函数的基础上估计的数据,和功能的主成分得分估计得到的条件步骤。这种条件估计方法在概念上是简单的,并且易于实现。一个关键的步骤是推导渐近一致性和分布结果在温和的条件下,使用工具从功能分析。对稀疏纵向数据的功能数据分析使得能够预测个体平滑轨迹,即使仅一个或几个测量值可用于受试者。渐近逐点和同时置信带预测的个人轨迹,渐近分布的基础上,同时带有限数量的组件的假设下。模型选择技术,如赤池信息准则,用于选择模型中的本征函数的数量对应的模型维数。该方法说明了一个模拟研究,纵向CD4数据的艾滋病患者的样本,和时间过程中的酵母细胞周期的基因表达数据。
We propose a nonparametric method to perform functional principal components analysis for the case of sparse longitudinal data. The method aims at irregularly spaced longitudinal data, where the number of repeated measurements available per subject is small. In contrast, classical functional data analysis requires a large number of regularly spaced measurements per subject. We assume that the repeated measurements are located randomly with a random number of repetitions for each subject and are determined by an underlying smooth random (subject-specific) trajectory plus measurement errors. Basic elements of our approach are the parsimonious estimation of the covariance structure and mean function of the trajectories, and the estimation of the variance of the measurement errors. The eigenfunction basis is estimated from the data, and functional principal components score estimates are obtained by a conditioning step. This conditional estimation method is conceptually simple and straightforward to implement. A key step is the derivation of asymptotic consistency and distribution results under mild conditions, using tools from functional analysis. Functional data analysis for sparse longitudinal data enables prediction of individual smooth trajectories even if only one or few measurements are available for a subject. Asymptotic pointwise and simultaneous confidence bands are obtained for predicted individual trajectories, based on asymptotic distributions, for simultaneous bands under the assumption of a finite number of components. Model selection techniques, such as the Akaike information criterion, are used to choose the model dimension corresponding to the number of eigenfunctions in the model. The methods are illustrated with a simulation study, longitudinal CD4 data for a sample of AIDS patients, and time-course gene expression data for the yeast cell cycle.