Predictive response-relevant clustering of expression data provides insights into disease processes.

Predictive response-relevant clustering of expression data provides insights into disease processes.
复制标题

DOI:
10.1093/nar/gkq550
复制
发表时间:
2010-11
影响因子:
14.9
通讯作者:
Dominiczak AF
Dominiczak AF
中科院分区:
生物学2区
文献类型:
--
作者:
Hopcroft LE;McBride MW;Harris KJ;Sampson AK;McClure JD;Graham D;Young G;Holyoake TL;Girolami MA;Dominiczak AF

文献摘要

参考文献

被引文献

相似文献

本文描述和说明了一种新的微阵列数据分析方法,耦合基于模型的聚类和二进制分类,形成集群的“响应相关”基因,也就是说,基因是信息时,区分不同的值的响应。预测随后使用适当的统计汇总每个基因簇,我们称之为“元协变量”表示的集群,在一个概率回归模型。我们首先通过分析白血病表达数据集来说明这种方法,然后密切关注盐敏感性高血压大鼠模型中肾脏基因表达数据集的元协变量分析。我们通过对这些数据的分析,探索了生物学方面的见解。特别是,我们确定了一个具有高度影响力的13个基因簇,包括三个转录因子(Arntl,Bhlhe 41和Npas 2),这是有牵连的,作为对高血压的保护,以应对增加饮食中的钠。使用独创性途径分析对该簇进行的功能和经典途径分析分别涉及转录激活和昼夜节律信号传导。虽然我们只使用表达式数据来说明我们的方法,但该方法适用于任何高维数据集。表达数据可在ArrayExpress(登录号E-MEXP-2514)获得,并且代码可在http://www.dcs.gla.ac.uk/inference/metacovariateanalysis/获得。
This article describes and illustrates a novel method of microarray data analysis that couples model-based clustering and binary classification to form clusters of `response-relevant' genes; that is, genes that are informative when discriminating between the different values of the response. Predictions are subsequently made using an appropriate statistical summary of each gene cluster, which we call the `meta-covariate' representation of the cluster, in a probit regression model. We first illustrate this method by analysing a leukaemia expression dataset, before focusing closely on the meta-covariate analysis of a renal gene expression dataset in a rat model of salt-sensitive hypertension. We explore the biological insights provided by our analysis of these data. In particular, we identify a highly influential cluster of 13 genes—including three transcription factors (Arntl, Bhlhe41 and Npas2)—that is implicated as being protective against hypertension in response to increased dietary sodium. Functional and canonical pathway analysis of this cluster using Ingenuity Pathway Analysis implicated transcriptional activation and circadian rhythm signalling, respectively. Although we illustrate our method using only expression data, the method is applicable to any high-dimensional datasets. Expression data are available at ArrayExpress (accession number E-MEXP-2514) and code is available at http://www.dcs.gla.ac.uk/inference/metacovariateanalysis/.
DOI: 10.4049/jimmunol.166.2.747
发表时间: 2001-01-15
影响因子: 4.4
作者:
Abe, R;Peng, T;Metz, CN
通讯作者: Metz, CN
DOI: 10.1126/science.286.5439.531
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Golub, TR;Slonim, DK;Lander, ES
通讯作者: Lander, ES
DOI: 10.1111/j.2517-6161.1995.tb02031.x
发表时间: 1995-01-01
影响因子: 5.8
作者:
BENJAMINI, Y;HOCHBERG, Y
通讯作者: HOCHBERG, Y
DOI: 10.1161/01.hyp.0000082766.63952.49
发表时间: 2003-08-01
期刊: HYPERTENSION
影响因子: 8.3
作者:
Mohri, T;Emoto, N;Yokoyama, M
通讯作者: Yokoyama, M
DOI: 10.1093/bioinformatics/bth419
发表时间: 2004-12-12
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Bae, K;Mallick, BK
通讯作者: Mallick, BK