Variable selection with false discovery rate control in deep neural networks

Variable selection with false discovery rate control in deep neural networks
复制标题

DOI:
10.1038/s42256-021-00308-z
复制
发表时间:
2019-09
影响因子:
23.8
通讯作者:
Zixuan Song;Jun Li
Zixuan Song;Jun Li
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zixuan Song;Jun Li

文献摘要

相似文献

深度神经网络以其高预测精度而闻名,但它们也以其黑箱性质和较差的可解释性而闻名。我们考虑了深度神经网络中的变量选择问题,即选择对输出具有显著预测能力的输入变量。现有的神经网络变量选择方法大多只适用于浅层网络或在大数据集上计算不可行,而且缺乏对所选变量质量的控制。在这里,我们提出了一种称为SurvNet的反向消除过程,它基于一种新的可变重要性度量,该度量适用于各种网络。更重要的是,SurvNet能够对选定变量的错误发现率进行经验估计和控制。此外,SurvNet自适应地确定每一步要消除多少变量,以便最大化选择效率。在不同的模拟和真实数据集上验证了SurvNet的有效性和效率,并与其他方法进行了性能比较。特别是,与基于仿冒的方法的系统比较表明,尽管它们对变量相关性较强的数据具有更严格的误发现率控制,但SurvNet通常具有更高的能力。
Deep neural networks are famous for their high prediction accuracy, but they are also known for their black-box nature and poor interpretability. We consider the problem of variable selection, that is, selecting the input variables that have significant predictive power on the output, in deep neural networks. Most existing variable selection methods for neural networks are only applicable to shallow networks or are computationally infeasible on large datasets; moreover, they lack a control on the quality of selected variables. Here we propose a backward elimination procedure called SurvNet, which is based on a new measure of variable importance that applies to a wide variety of networks. More importantly, SurvNet is able to estimate and control the false discovery rate of selected variables empirically. Further, SurvNet adaptively determines how many variables to eliminate at each step in order to maximize the selection efficiency. The validity and efficiency of SurvNet are shown on various simulated and real datasets, and its performance is compared with other methods. Especially, a systematic comparison with knockoff-based methods shows that although they have more rigorous false discovery rate control on data with strong variable correlation, SurvNet usually has higher power.