An empirical bayes adjustment to increase the sensitivity of detecting differentially expressed genes in microarray experiments

An empirical bayes adjustment to increase the sensitivity of detecting differentially expressed genes in microarray experiments
复制标题

DOI:
10.1093/bioinformatics/btg396
复制
发表时间:
2004-01-22
期刊:
影响因子:
5.8
通讯作者:
Datta, S
Datta, S
中科院分区:
生物学3区
文献类型:
--
作者:
Datta, S;Satten, GA;Datta, S

文献摘要

被引文献

相似文献

动机:检测差异表达基因是微阵列实验的主要目标之一。在没有控制总体(实验)1型错误率的情况下,对每个基因进行配对比较是不合适的。Dudoit等人。主张使用基于排列的递减P值调整来校正个体(即每个基因)两个样本t检验的观察到的显著水平。结果:在本文中,我们考虑了对应于多个组织类型的基因表达水平的方差分析公式。我们提供基于重采样的逐级调整,以校正每个基因和每对组织类型比较的单个ANOVA t检验的观察到的显著水平。更重要的是,我们引入了一种新的经验贝叶斯调整t-检验统计量,可以合并到逐步下降过程中。使用模拟数据,我们表明,经验贝叶斯调整提高了检测差异表达基因的灵敏度高达16%,同时保持了高水平的特异性。这一调整还在一定程度上降低了错误未发现率,但代价是错误发现率略有增加。我们使用由正常细胞、腺瘤细胞和癌细胞的寡核苷酸阵列组成的人类结肠癌数据集来说明我们的方法。在比较正常细胞和腺瘤细胞时,差异表达水平的基因数约为50个,而在比较腺瘤和癌细胞时,差异表达水平的基因数约为5个。这份清单包括以前已知的与结肠癌相关的基因以及一些新的基因。
Motivation: Detection of differentially expressed genes is one of the major goals of microarray experiments. Pairwise comparison for each gene is not appropriate without controlling the overall (experimentwise) type 1 error rate. Dudoit et al. have advocated use of permutation-based step-down P-value adjustments to correct the observed significance levels for the individual (i.e. for each gene) two sample t-tests.Results: In this paper, we consider an ANOVA formulation of the gene expression levels corresponding to multiple tissue types. We provide resampling-based step-down adjustments to correct the observed significance levels for the individual ANOVA t-tests for each gene and for each pair of tissue type comparisons. More importantly, we introduce a novel empirical Bayes adjustment to the t-test statistics that can be incorporated into the step-down procedure. Using simulated data, we show that the empirical Bayes adjustment improved the sensitivity of detecting differentially expressed genes up to 16%, while maintaining a high level of specificity. This adjustment also reduces the false non-discovery rate to some degree at the cost of a modest increase in the false discovery rate. We illustrate our approach using a human colon cancer dataset consisting of oligonucleotide arrays of normal, adenoma and carcinoma cells. The number of genes with differential expression level declared statistically significant was about 50 when comparing normal to adenoma cells and about five when comparing adenoma to carcinoma cells. This list includes genes previously known to be associated with colon cancer as well as some novel genes.