False Discovery Rate Control via Data Splitting

False Discovery Rate Control via Data Splitting
复制标题

DOI:
10.1080/01621459.2022.2060113
复制
发表时间:
2022-05-25
影响因子:
3.7
通讯作者:
Liu, Jun S.
Liu, Jun S.
中科院分区:
数学1区
文献类型:
--
作者:
Dai, Chenguang;Lin, Buyu;Liu, Jun S.

文献摘要

被引文献

相似文献

选择与给定响应变量相关联的相关特征是许多科学领域中的重要问题。通过错误发现率(FDR)控制来量化选择结果的质量和不确定性最近受到关注。本文介绍了一种数据分裂的方法(简称为“DS”),渐近控制FDR,同时保持高功率。对于每个特征,DS通过数据分裂估计两个独立的回归系数来构造检验统计量。FDR控制是通过利用统计量的属性来实现的,对于任何空特征,其采样分布关于零对称;而对于相关特征,其采样分布具有正均值。此外,提出了一种多数据分裂(MDS)方法,以稳定选择结果,提高功率。令人惊讶的是,在FDR处于控制之下的情况下,MDS不仅有助于克服由数据分裂引起的功率损失,而且与所考虑的所有其他方法相比,还导致更低的错误发现比例(FDP)的方差。大量的仿真研究和实际数据的应用表明,所提出的方法是强大的未知分布的特征,易于实现和计算效率,往往是最强大的竞争对手,特别是当信号是弱的,相关性或偏相关性的特征是高。本文的补充材料可在网上查阅。
Selecting relevant features associated with a given response variable is an important problem in many scientific fields. Quantifying quality and uncertainty of a selection result via false discovery rate (FDR) control has been of recent interest. This article introduces a data-splitting method (referred to as "DS") to asymptotically control the FDR while maintaining a high power. For each feature, DS constructs a test statistic by estimating two independent regression coefficients via data splitting. FDR control is achieved by taking advantage of the statistic's property that, for any null feature, its sampling distribution is symmetric about zero; whereas for a relevant feature, its sampling distribution has a positive mean. Furthermore, a Multiple Data Splitting (MDS) method is proposed to stabilize the selection result and boost the power. Surprisingly, with the FDR under control, MDS not only helps overcome the power loss caused by data splitting, but also results in a lower variance of the false discovery proportion (FDP) compared with all other methods in consideration. Extensive simulation studies and a real-data application show that the proposed methods are robust to the unknown distribution of features, easy to implement and computationally efficient, and are often the most powerful ones among competitors especially when the signals are weak and correlations or partial correlations among features are high. Supplementary materials for this article are available online.