A Dirichlet process mixture model for clustering longitudinal gene expression data.

A Dirichlet process mixture model for clustering longitudinal gene expression data.
复制标题

DOI:
10.1002/sim.7374
复制
发表时间:
2017-09-30
影响因子:
2
通讯作者:
Warren JL
Warren JL
中科院分区:
医学3区
文献类型:
--
作者:
Sun J;Herazo-Maya JD;Kaminski N;Zhao H;Warren JL

文献摘要

参考文献

被引文献

相似文献

子群识别(聚类)是生物医学研究中的重要问题。基因表达谱通常用于定义亚群。纵向基因表达谱可能比单独的基线谱提供关于疾病进展的额外信息。因此,借助纵向基因表达数据进行亚群鉴定可以更加准确和有效。然而,现有的统计方法无法充分利用这些数据进行患者聚类。在本文中,我们介绍了一种新的基于纵向基因表达谱的贝叶斯聚类方法。该方法称为BClustLonG,采用线性混合效应框架对基因随时间的轨迹进行建模,并根据所有基因的回归系数共同进行聚类。为了解释基因间的相关性,减轻高维度的挑战,我们对回归系数采用因子分析模型。利用Dirichlet过程先验分布作为回归系数的均值来诱导聚类。通过广泛的仿真研究,我们表明bclusterlong比其他聚类方法具有更高的性能。当应用于严重受伤(烧伤或创伤)患者的数据集时,我们的模型能够识别出有趣的亚组。
Subgroup identification (clustering) is an important problem in biomedical research. Gene expression profiles are commonly utilized to define subgroups. Longitudinal gene expression profiles might provide additional information on disease progression than what is captured by baseline profiles alone. Therefore, subgroup identification could be more accurate and effective with the aid of longitudinal gene expression data. However, existing statistical methods are unable to fully utilize these data for patient clustering. In this article, we introduce a novel clustering method in the Bayesian setting based on longitudinal gene expression profiles. This method, called BClustLonG, adopts a linear mixed-effects framework to model the trajectory of genes over time while clustering is jointly conducted based on the regression coefficients obtained from all genes. In order to account for the correlations among genes and alleviate the high dimensionality challenges, we adopt a factor analysis model for the regression coefficients. The Dirichlet process prior distribution is utilized for the means of the regression coefficients to induce clustering. Through extensive simulation studies, we show that BClustLonG has improved performance over other clustering methods. When applied to a dataset of severely injured (burn or trauma) patients, our model is able to identify interesting subgroups.
DOI: 10.1198/016214505000000187
发表时间: 2006-03-01
影响因子: 3.7
作者:
Heard, NA;Holmes, CC;Stephens, DA
通讯作者: Stephens, DA
DOI: 10.1214/12-aoas580
发表时间: 2013-03-01
影响因子: 1.8
作者:
Komarek, Arnost;Komarkova, Lenka
通讯作者: Komarkova, Lenka
DOI: 10.1093/bioinformatics/bth068
发表时间: 2004-05-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Medvedovic, M;Yeung, KY;Bumgarner, RE
通讯作者: Bumgarner, RE
DOI: 10.1038/35000501
发表时间: 2000-02-03
期刊: NATURE
影响因子: 64.8
作者:
Alizadeh, AA;Eisen, MB;Staudt, LM
通讯作者: Staudt, LM
DOI: 10.1038/nrneurol.2013.278
发表时间: 2014-02
期刊: Nature reviews. Neurology
影响因子: --
作者:
通讯作者: --