Comparison of Bayesian predictive methods for model selection

Comparison of Bayesian predictive methods for model selection
复制标题

DOI:
10.1007/s11222-016-9649-y
复制
发表时间:
2017-05-01
影响因子:
2.2
通讯作者:
Vehtari, Aki
Vehtari, Aki
中科院分区:
数学2区
文献类型:
--
作者:
Piironen, Juho;Vehtari, Aki

文献摘要

被引文献

相似文献

本文的目的是比较几个广泛使用的贝叶斯模型选择方法在实际的模型选择问题,突出他们的差异,并给出建议的首选方法。我们专注于回归和分类的变量子集选择,并使用模拟和真实的世界数据进行了几个数值实验。结果表明,优化效用估计,如交叉验证(CV)得分是容易找到过拟合模型,由于相对较高的方差时,数据是稀缺的效用估计。这也可能导致在对所选模型的性能评估中产生大量选择诱导的偏见和乐观。从预测的角度来看,最好的结果是通过考虑模型的不确定性,形成完整的包容性模型,如贝叶斯模型平均解决方案的候选模型。如果包含模型过于复杂,则可以通过投影方法对其进行鲁棒简化,在该方法中,完整模型的信息被投影到子模型上。这种方法比基于CV分数的选择更不容易过拟合。总体而言,投影方法似乎也优于最大后验模型和最可能的变量的选择。该研究还表明,模型选择可以大大受益于使用搜索过程之外的交叉验证,用于指导模型大小的选择和评估最终选择的模型的预测性能。
The goal of this paper is to compare several widely used Bayesian model selection methods in practical model selection problems, highlight their differences and give recommendations about the preferred approaches. We focus on the variable subset selection for regression and classification and perform several numerical experiments using both simulated and real world data. The results show that the optimization of a utility estimate such as the cross-validation (CV) score is liable to finding overfitted models due to relatively high variance in the utility estimates when the data is scarce. This can also lead to substantial selection induced bias and optimism in the performance evaluation for the selected model. From a predictive viewpoint, best results are obtained by accounting for model uncertainty by forming the full encompassing model, such as the Bayesian model averaging solution over the candidate models. If the encompassing model is too complex, it can be robustly simplified by the projection method, in which the information of the full model is projected onto the submodels. This approach is substantially less prone to overfitting than selection based on CV-score. Overall, the projection method appears to outperform also the maximum a posteriori model and the selection of the most probable variables. The study also demonstrates that the model selection can greatly benefit from using cross-validation outside the searching process both for guiding the model size selection and assessing the predictive performance of the finally selected model.