VARIABLE SELECTION FOR CLASSIFICATION WITH DERIVATIVE-INDUCED REGULARIZATION

VARIABLE SELECTION FOR CLASSIFICATION WITH DERIVATIVE-INDUCED REGULARIZATION
复制标题

通过导数诱导正则化进行分类的变量选择

DOI:
10.5705/ss.202018.0086
复制
发表时间:
2020-10-01
期刊:
影响因子:
1.4
通讯作者:
Wang, Junhui
Wang, Junhui
中科院分区:
数学3区
文献类型:
--
作者:
He, Xin;Lv, Shaogao;Wang, Junhui

文献摘要

被引文献

相似文献

尽管过去二十年来对变量选择进行了广泛的研究,但关于分类变量选择的研究很少,特别是在没有对模型做出任何假设的情况下。在本文中,我们提出了一个通用的变量选择框架,通过考察条件概率进行分类。该框架利用具有导数诱导稀疏性的支持向量机来描述,它不需要明确的模型假设,并且充分利用了再生核Hilbert空间的数学特性。与已有的许多方法相比,我们提出的方法引入了一个凸优化问题,并充分利用了光滑RKHS中梯度的再生性,充分利用了梯度信息。该方法也可以看作是经典支持向量机的推广,在稀疏分类中取得了较好的经验性能。重要的是,建立了该方法的估计相合性和子集选择性质。最后,验证了《统计学报:预印本DOI:10.5705/ss.202018.0086》方法的有效性
Despite extensive research on variable selection over the past two decades, few studies exist on variable selection for classification, particularly when no assumptions are made about the model. In this paper, we propose a general variable selection framework for classification by examining the conditional probability. The proposed framework is illustrated by means of support vector machine (SVM) with derivative-induced sparsity, which makes no explicit model assumption, and takes full advantage of the mathematical properties of the reproducing kernel Hilbert space (RKHS). In contrast to many existing methods, our proposed method leads to a convex optimization task, and fully exploits gradient information by using the reproducing property of gradients in smooth RKHSs. The proposed method can also be viewed as a generalization of the classical SVM, and achieves superior empirical performance in sparse classification. Importantly, the estimation consistency and subset selection properties of the proposed method are established. Lastly, the effectiveness is of the method Statistica Sinica: Preprint doi:10.5705/ss.202018.0086