False Discovery Rate Estimation for Stability Selection: Application to Genome-Wide Association Studies

False Discovery Rate Estimation for Stability Selection: Application to Genome-Wide Association Studies
复制标题

DOI:
10.2202/1544-6115.1663
复制
发表时间:
2011-01-01
影响因子:
0.9
通讯作者:
Richardson, Sylvia
Richardson, Sylvia
中科院分区:
数学4区
文献类型:
--
作者:
Ahmed, Ismail;Hartikainen, Anna-Liisa;Richardson, Sylvia

文献摘要

被引文献

相似文献

稳定性选择将惩罚回归与二次采样相结合,是一种很有前景的超高维变量选择算法。这项工作的动机是在全基因组关联研究(GWAS)的背景下进行评估。其使用的一个关键方面在于选择一种决策规则,该规则能够解释所实现的大量比较。当前的决策规则依赖于通过理论上推导的上限来控制族明智错误率(FWER)。或者,我们建议根据更自由的错误发现率(FDR)标准来设置检测阈值。我们提出的估计过程依赖于排列。该过程根据模拟遗传数据的各种相关结构的几种场景进行模拟评估,并与原始 FWER 上限进行比较。所提出的过程被证明不太保守,并且能够拾取比 FWER 上限更多的真实信号。最后,通过对芬兰北部出生队列中脂质表型(高密度脂蛋白,HDL)的 GWAS 分析说明了所提出的方法。
Stability Selection, which combines penalized regression with subsampling, is a promising algorithm to perform variable selection in ultra high dimension. This work is motivated by its evaluation in the context of genome-wide association studies (GWAS). One critical aspect for its use lies in the choice of a decision rule that accounts for the massive number of comparisons realised. The current decision rule relies on the control of the Family Wise Error Rate (FWER) by means of an upper bound derived theoretically. Alternatively, we propose to set the detection threshold according to the more liberal false discovery rate (FDR) criterion. The procedure we propose for its estimation relies on permutations. This procedure is evaluated by simulations according to several scenarios mimicking various correlation structures of genetic data and is compared to the original FWER upper bound. The proposed procedure is shown to be less conservative, and able to pick up more true signals than the FWER upper bound. Finally, the proposed methodology is illustrated on a GWAS analysis of a lipid phenotype (high-density lipoproteins, HDL) in the Northern Finland Birth Cohort.