ZAP: Z -Value Adaptive Procedures for False Discovery Rate Control with Side Information

ZAP: Z -Value Adaptive Procedures for False Discovery Rate Control with Side Information
复制标题

ZAP:利用辅助信息进行错误发现率控制的 Z 值自适应程序

DOI:
10.1111/rssb.12557
复制
发表时间:
2022
期刊:
Journal of the Royal Statistical Society Series B: Statistical Methodology
影响因子:
--
通讯作者:
Sun, Wenguang
Sun, Wenguang
中科院分区:
--
文献类型:
--
作者:
Leung, Dennis;Sun, Wenguang

文献摘要

相似文献

协变量的自适应多重测试是近年来引起广泛关注的一个重要研究方向。人们广泛认识到,利用辅助协变量提供的辅助信息可以提高错误发现率 (FDR) 程序的能力。目前,大多数此类程序都是以 p 值作为主要统计数据来设计的。然而,对于双向假设,将主要统计数据(称为 asz 值、intop 值)转换的常用数据处理步骤不仅会导致主要统计数据所携带的信息丢失,而且还会破坏协变量协助 FDR 推理的能力。我们开发了基于 z 值的协变量自适应(ZAP)方法,该方法对由 z 值和协变量联合编码的完整结构信息进行操作。它试图通过工作模型模拟 oraclez 值过程,其拒绝区域与 p 值自适应测试方法的拒绝区域显着不同。 ZAP 的关键优势在于,即使工作模型指定错误,也可以通过最少的假设来保证 FDR 控制。我们使用模拟和实际数据证明了 ZAP 的最先进性能,这表明与基于 p 值的方法相比,效率增益可以是可观的。我们的方法是在 R 包zap 中实现的。
Adaptive multiple testing with covariates is an important research direction that has gained major attention in recent years. It has been widely recognised that leveraging side information provided by auxiliary covariates can improve the power of false discovery rate (FDR) procedures. Currently, most such procedures are devised withp‐values as their main statistics. However, for two‐sided hypotheses, the usual data processing step that transforms the primary statistics, known asz‐values, intop‐values not only leads to a loss of information carried by the main statistics, but can also undermine the ability of the covariates to assist with the FDR inference. We develop az‐value based covariate‐adaptive (ZAP) methodology that operates on the intact structural information encoded jointly by thez‐values and covariates. It seeks to emulate the oraclez‐value procedure via a working model, and its rejection regions significantly depart from those of thep‐value adaptive testing approaches. The key strength of ZAP is that the FDR control is guaranteed with minimal assumptions, even when the working model is misspecified. We demonstrate the state‐of‐the‐art performance of ZAP using both simulated and real data, which shows that the efficiency gain can be substantial in comparison withp‐value‐based methods. Our methodology is implemented in the R packagezap.