Re-sampling strategy to improve the estimation of number of null hypotheses in FDR control under strong correlation structures

Re-sampling strategy to improve the estimation of number of null hypotheses in FDR control under strong correlation structures
复制标题

DOI:
10.1186/1471-2105-8-157
复制
发表时间:
2007-05-18
期刊:
影响因子:
3
通讯作者:
Perkins, David L.
Perkins, David L.
中科院分区:
生物学4区
文献类型:
--
作者:
Lu, Xin;Perkins, David L.

文献摘要

被引文献

相似文献

背景:在进行多重假设检验时,重要的是要控制假阳性的数量,或错误发现率(FDR)。然而,在控制FDR和最大化功率之间存在折衷。已经提出了几种方法,如q值法,以估计真零假设在检验假设中所占的比例,并将此估计用于FDR的控制。这些方法通常依赖于假设检验统计量是独立的(或仅弱相关)。然而,许多类型的数据,例如微阵列数据,通常包含大规模的相关性结构。我们的目标是开发方法来控制FDR,同时保持一个更大的水平的权力,在高度相关的datasets通过提高估计的零假设的比例。结果:我们发现,当强相关性存在于数据之间,这是常见的微阵列数据集,估计的零假设的比例可能是高度可变的,导致在FDR的变化的高水平。因此,我们开发了一种重采样策略,通过打破基因表达值之间的相关性来减少变异,然后使用选择重采样估计值的上四分位数的保守策略来获得FDR的强控制。通过对实际微阵列数据集的模拟研究和扰动,与q值等竞争方法相比,对零假设的比例产生轻微偏差的估计,但具有较低的均方误差。当选择控制相同FDR水平的基因时,我们的方法平均具有显著较低的错误发现率,以换取功率的轻微降低。
Background: When conducting multiple hypothesis tests, it is important to control the number of false positives, or the False Discovery Rate (FDR). However, there is a tradeoff between controlling FDR and maximizing power. Several methods have been proposed, such as the q-value method, to estimate the proportion of true null hypothesis among the tested hypotheses, and use this estimation in the control of FDR. These methods usually depend on the assumption that the test statistics are independent (or only weakly correlated). However, many types of data, for example microarray data, often contain large scale correlation structures. Our objective was to develop methods to control the FDR while maintaining a greater level of power in highly correlated datasets by improving the estimation of the proportion of null hypotheses.Results: We showed that when strong correlation exists among the data, which is common in microarray datasets, the estimation of the proportion of null hypotheses could be highly variable resulting in a high level of variation in the FDR. Therefore, we developed a re-sampling strategy to reduce the variation by breaking the correlations between gene expression values, then using a conservative strategy of selecting the upper quartile of the re-sampling estimations to obtain a strong control of FDR.Conclusion: With simulation studies and perturbations on actual microarray datasets, our method, compared to competing methods such as q-value, generated slightly biased estimates on the proportion of null hypotheses but with lower mean square errors. When selecting genes with controlling the same FDR level, our methods have on average a significantly lower false discovery rate in exchange for a minor reduction in the power.