Best subset selection via cross-validation criterion

Best subset selection via cross-validation criterion
复制标题

DOI:
10.1007/s11750-020-00538-1
复制
发表时间:
2020-07
期刊:
TOP
影响因子:
1.7
通讯作者:
Yuichi Takano;Ryuhei Miyashiro
Yuichi Takano;Ryuhei Miyashiro
中科院分区:
管理学4区
文献类型:
--
作者:
Yuichi Takano;Ryuhei Miyashiro

文献摘要

相似文献

本文讨论了线性回归模型中选择最佳解释变量子集的交叉验证准则。与使用统计标准(例如,Mallows的,Akaike信息准则和贝叶斯信息准则),交叉验证只需要温和的假设,即样本是相同分布的,训练和验证样本是独立的。出于这个原因,交叉验证标准预计在大多数涉及预测方法的情况下都能很好地工作。本文的目的是建立一个混合整数优化方法来选择最佳的解释变量子集,通过交叉验证准则。这个子集选择问题可以用公式表示为一个双层MIO问题。然后,我们把它归结为一个单级混合整数二次优化问题,这可以通过使用优化软件精确求解。我们的方法的有效性进行了评估,通过模拟实验,通过比较与基于启发式准则的穷举搜索算法和正则化回归。我们的仿真结果表明,当信噪比较低时,我们的方法对于子集选择和预测都具有良好的准确性。
This paper is concerned with the cross-validation criterion for selecting the best subset of explanatory variables in a linear regression model. In contrast with the use of statistical criteria (e.g., Mallows’, the Akaike information criterion, and the Bayesian information criterion), cross-validation requires only mild assumptions, namely, that samples are identically distributed and that training and validation samples are independent. For this reason, the cross-validation criterion is expected to work well in most situations involving predictive methods. The purpose of this paper is to establish a mixed-integer optimization approach to selecting the best subset of explanatory variables via the cross-validation criterion. This subset-selection problem can be formulated as a bilevel MIO problem. We then reduce it to a single-level mixed-integer quadratic optimization problem, which can be solved exactly by using optimization software. The efficacy of our method is evaluated through simulation experiments by comparison with statistical-criterion-based exhaustive search algorithms and-regularized regression. Our simulation results demonstrate that, when the signal-to-noise ratio was low, our method delivered good accuracy for both subset selection and prediction.