CONTRA: Contrarian statistics for controlled variable selection

CONTRA: Contrarian statistics for controlled variable selection
复制标题

DOI:
--
复制
发表时间:
2021-04
期刊:
Proceedings of machine learning research
影响因子:
--
通讯作者:
Mukund Sudarshan;A. Puli;Lakshminarayanan Subramanian;S. Sankararaman;R. Ranganath
Mukund Sudarshan;A. Puli;Lakshminarayanan Subramanian;S. Sankararaman;R. Ranganath
中科院分区:
其他
文献类型:
--
作者:
Mukund Sudarshan;A. Puli;Lakshminarayanan Subramanian;S. Sankararaman;R. Ranganath

文献摘要

相似文献

保持随机化检验(HRT)发现了一组最能预测反应的协变量。给定协变量分布,HRT可以显式控制错误发现率(FDR)。然而,如果这种分布是未知的,必须从数据中估计,HRT可以膨胀FDR。为了缓解FDR的膨胀,我们提出了反向随机化检验(CONTRA),该检验是专门针对必须根据数据估计协变量分布甚至可能被错误指定的情况而设计的。我们的关键见解是使用两个“逆向”概率模型的平等混合来确定协变量的重要性。一个模型用真实的数据拟合,而另一个模型用相同的数据拟合,但是用来自协变量分布估计值的样本替换了被检验的协变量。CONTRA具有足够的灵活性,可以渐进地实现1的幂,当协变量分布被错误指定时,与最先进的CVS方法相比可以降低FDR,并且在高维和大样本量中计算效率高。我们进一步证明了CONTRA在众多合成基准上的有效性,并强调了其在遗传数据集上的能力。
The holdout randomization test (HRT) discovers a set of covariates most predictive of a response. Given the covariate distribution, HRTs can explicitly control the false discovery rate (FDR). However, if this distribution is unknown and must be estimated from data, HRTs can inflate the FDR. To alleviate the inflation of FDR, we propose the contrarian randomization test (CONTRA), which is designed explicitly for scenarios where the covariate distribution must be estimated from data and may even be misspecified. Our key insight is to use an equal mixture of two "contrarian" probabilistic models in determining the importance of a covariate. One model is fit with the real data, while the other is fit using the same data, but with the covariate being tested replaced with samples from an estimate of the covariate distribution. CONTRA is flexible enough to achieve a power of 1 asymptotically, can reduce the FDR compared to state-of-the-art CVS methods when the covariate distribution is misspecified, and is computationally efficient in high dimensions and large sample sizes. We further demonstrate the effectiveness of CONTRA on numerous synthetic benchmarks, and highlight its capabilities on a genetic dataset.