CLUSTERING FOR MULTIVARIATE CONTINUOUS AND DISCRETE LONGITUDINAL DATA

CLUSTERING FOR MULTIVARIATE CONTINUOUS AND DISCRETE LONGITUDINAL DATA
复制标题

DOI:
10.1214/12-aoas580
复制
发表时间:
2013-03-01
影响因子:
1.8
通讯作者:
Komarkova, Lenka
Komarkova, Lenka
中科院分区:
数学4区
文献类型:
--
作者:
Komarek, Arnost;Komarkova, Lenka

文献摘要

被引文献

相似文献

在纵向研究和常规临床随访中,通常会收集受试者的连续和离散的多个结果。为了激励我们的工作,我们考虑了一项对原发性胆汁性肝硬化(PBC)患者的纵向研究,这些患者具有连续的胆红素水平,离散的血小板计数和血管畸形的二分类指征作为此类纵向结果的例子。一个明显的要求是使用所有的结果值对受试者进行分组(例如,在临床环境中具有相似预后的受试者组)。近年来,人们提出了许多基于纵向(或其他相关)结果的分类方法,不仅针对生物统计学等传统领域,还针对快速发展的生物信息学和许多其他领域。然而,大多数可用的方法只考虑连续结果作为分类的基础,或者如果考虑非连续结果,则不与其他不同性质的结果结合使用。在这里,我们提出了一种统计方法,在重复测量几个不同性质的纵向结果的基础上,将具有先验未知特征的受试者聚类(分类)到预先指定数量的组中。该方法依赖于经典广义线性混合模型的多元扩展,其中额外假设随机效应的混合分布。我们基于模型的贝叶斯规范和基于模拟的马尔可夫链蒙特卡罗方法进行推理。为了在实践中应用该方法,我们准备了在R中使用的现成软件(http://www.R-project.org)。我们还讨论了分类中不确定性的评估,并讨论了最近提出的模型比较方法的使用-在我们的案例中,基于Plummer [Biostatistics 9(2008) 523-539]提出的惩罚后验偏差选择一些聚类。
Multiple outcomes, both continuous and discrete, are routinely gathered on subjects in longitudinal studies and during routine clinical follow-up in general. To motivate our work, we consider a longitudinal study on patients with primary biliary cirrhosis (PBC) with a continuous bilirubin level, a discrete platelet count and a dichotomous indication of blood vessel malformations as examples of such longitudinal outcomes. An apparent requirement is to use all the outcome values to classify the subjects into groups (e. g., groups of subjects with a similar prognosis in a clinical setting). In recent years, numerous approaches have been suggested for classification based on longitudinal (or otherwise correlated) outcomes, targeting not only traditional areas like biostatistics, but also rapidly evolving bioinformatics and many others. However, most available approaches consider only continuous outcomes as a basis for classification, or if noncontinuous outcomes are considered, then not in combination with other outcomes of a different nature. Here, we propose a statistical method for clustering (classification) of subjects into a prespecified number of groups with a priori unknown characteristics on the basis of repeated measurements of several longitudinal outcomes of a different nature. This method relies on a multivariate extension of the classical generalized linear mixed model where a mixture distribution is additionally assumed for random effects. We base the inference on a Bayesian specification of the model and simulation-based Markov chain Monte Carlo methodology. To apply the method in practice, we have prepared ready-to-use software for use in R (http://www.R-project.org). We also discuss evaluation of uncertainty in the classification and also discuss usage of a recently proposed methodology for model comparison-the selection of a number of clusters in our case-based on the penalized posterior deviance proposed by Plummer [Biostatistics 9 (2008) 523-539].