Surrogate variable analysis using partial least squares (SVA-PLS) in gene expression studies

Surrogate variable analysis using partial least squares (SVA-PLS) in gene expression studies
复制标题

DOI:
10.1093/bioinformatics/bts022
复制
发表时间:
2012-03-15
期刊:
影响因子:
5.8
通讯作者:
Datta, Susmita
Datta, Susmita
中科院分区:
生物学3区
文献类型:
--
作者:
Chakraborty, Sutirtha;Datta, Somnath;Datta, Susmita

文献摘要

被引文献

相似文献

动机:在典型的基因表达谱研究中,我们的主要目标是鉴定两种不同组织类型样品之间差异表达的基因。通常,实施标准方差分析(ANOVA)/回归以从它们各自的表达水平阵列鉴定这些基因对两种类型的样品的相对影响。但是,当这些阵列中存在未知的可变性来源(归因于不同生物,环境或其他相关因素的潜在变量)时,这种技术就存在根本性缺陷。这些因素扭曲了两种组织类型之间差异基因表达的真实情况,并引入了表达异质性的虚假信号。结果,许多实际上差异表达的基因没有被检测到,而许多其他基因被错误地鉴定为阳性。此外,这些扭曲对于不同的基因可能是不同的。因此,也不可能通过简单的数组规范化来消除这些变化。这种双向错误可能导致灵敏度和特异性的严重损失,从而导致潜在的多重测试问题的严重效率低下。在这项工作中,我们试图通过偏最小二乘法(PLS)识别潜在因素在基因表达谱研究中的隐藏效应,并将PLS识别的这些隐藏效应的特征作为协变量应用ANCOVA技术,以识别两种相关组织类型之间真正差异表达的基因。我们比较我们的方法SVA-PLS与标准方差分析和一个相对较新的技术的替代变量分析(SVA)的性能,在各种各样的模拟设置(将不同的影响的隐藏变量,在不同的信号强度和基因分组的情况下)。在所有设置中,我们的方法产生最高的灵敏度,同时保持相对合理的值的特异性,错误的发现率和错误的非发现率。我们的方法应用于急性巨核细胞白血病的基因表达谱分析表明,我们的方法检测到额外的6个基因,这是错过了标准方差分析方法以及SVA,但可能与这种疾病,可以看出从挖掘现有的文献。
Motivation: In a typical gene expression profiling study, our prime objective is to identify the genes that are differentially expressed between the samples from two different tissue types. Commonly, standard analysis of variance (ANOVA)/regression is implemented to identify the relative effects of these genes over the two types of samples from their respective arrays of expression levels. But, this technique becomes fundamentally flawed when there are unaccounted sources of variability in these arrays (latent variables attributable to different biological, environmental or other factors relevant in the context). These factors distort the true picture of differential gene expression between the two tissue types and introduce spurious signals of expression heterogeneity. As a result, many genes which are actually differentially expressed are not detected, whereas many others are falsely identified as positives. Moreover, these distortions can be different for different genes. Thus, it is also not possible to get rid of these variations by simple array normalizations. This both-way error can lead to a serious loss in sensitivity and specificity, thereby causing a severe inefficiency in the underlying multiple testing problem. In this work, we attempt to identify the hidden effects of the underlying latent factors in a gene expression profiling study by partial least squares (PLS) and apply ANCOVA technique with the PLS-identified signatures of these hidden effects as covariates, in order to identify the genes that are truly differentially expressed between the two concerned tissue types.Results: We compare the performance of our method SVA-PLS with standard ANOVA and a relatively recent technique of surrogate variable analysis (SVA), on a wide variety of simulation settings (incorporating different effects of the hidden variable, under situations with varying signal intensities and gene groupings). In all settings, our method yields the highest sensitivity while maintaining relatively reasonable values for the specificity, false discovery rate and false non-discovery rate. Application of our method to gene expression profiling for acute megakaryoblastic leukemia shows that our method detects an additional six genes, that are missed by both the standard ANOVA method as well as SVA, but may be relevant to this disease, as can be seen from mining the existing literature.