Identifying differentially expressed genes using false discovery rate controlling procedures

Identifying differentially expressed genes using false discovery rate controlling procedures
复制标题

DOI:
10.1093/bioinformatics/btf877
复制
发表时间:
2003-02-12
期刊:
影响因子:
5.8
通讯作者:
Benjamini, Y
Benjamini, Y
中科院分区:
生物学3区
文献类型:
--
作者:
Reiner, A;Yekutieli, D;Benjamini, Y

文献摘要

被引文献

相似文献

动机:DNA微阵列最近被用于同时监测数千个基因的表达水平,并识别那些差异表达的基因。当测试的基因数量变大时,错误识别(第一类错误)的概率可能会急剧增加。归因于基因共同调节的测试统计数据与基因表达水平测量误差的依赖性之间的相关性使问题进一步复杂化。在本文中,我们通过采用错误发现率(FDR)控制方法来解决这个非常大的多重性问题。为了解决相关性问题,我们提出了三种基于重抽样的FDR控制程序,它们解释了检验统计分布,并将它们的性能与Benjamini和Hochberg(1995)中简单应用的线性递增程序的性能进行了比较。结果:对比仿真分析表明,四种差错率控制程序均能将差错率控制在期望的水平,并且比家族式差错率控制程序保留了更多的功率。在功率方面,使用每个测试统计数据的边际分布的重采样显著提高了相对于幼稚的性能。以更复杂的算法为代价,通过对测试统计数据的联合分布进行重采样并估计FDR控制水平的基于重采样的程序,可以实现最高功率。
Motivation: DNA microarrays have recently been used for the purpose of monitoring expression levels of thousands of genes simultaneously and identifying those genes that are differentially expressed. The probability that a false identification (type I error) is committed can increase sharply when the number of tested genes gets large. Correlation between the test statistics attributed to gene co-regulation and dependency in the measurement errors of the gene expression levels further complicates the problem. In this paper we address this very large multiplicity problem by adopting the false discovery rate (FDR) controlling approach. In order to address the dependency problem, we present three resampling-based FDR controlling procedures, that account for the test statistics distribution, and compare their performance to that of the naive application of the linear step-up procedure in Benjamini and Hochberg (1995). The procedures are studied using simulated microarray data, and their performance is examined relative to their ease of implementation.Results: Comparative simulation analysis shows that all four FDR controlling procedures control the FDR at the desired level, and retain substantially more power then the family-wise error rate controlling procedures. In terms of power, using resampling of the marginal distribution of each test statistics substantially improves the performance over the naive one. The highest power is achieved, at the expense of a more sophisticated algorithm, by the resampling-based procedures that resample the joint distribution of the test statistics and estimate the level of FDR control.