Differential Description Length for Hyperparameter Selection in Supervised Learning
Differential Description Length for Hyperparameter Selection in Supervised Learning
复制标题
DOI:
--
复制
发表时间:
2020-10
期刊:
影响因子:
--
通讯作者:
M. Abolfazli;A. Høst-Madsen;June Zhang
中科院分区:
文献类型:
--
作者:
M. Abolfazli;A. Høst-Madsen;June Zhang
Minimum description length (MDL) is an established method for model selection. For supervised learning problems, cross-validation is often used for model selection in practice. Reasons are 1) MDL is difficult to apply directly to data; 2) MDL may make restrictive statistical assumptions that decrease performance; and 3) MDL does not directly aim to minimize generalization error. In this paper, we introduce a modification to MDL, which we call differential description length (DDL). DDL partitions the data so that the codelength(s) it computes, reflects the conditional probability of seeing ‘new’ data given ‘old’ data. This differential codelength is what allows DDL to estimate generalization error like cross-validation. DDL is also better than cross-validation because it allows the learning algorithm to use the entire data without having to withhold subsets for validation and testing. Compared with MDL, DDL has both better performance (in finding models with smaller generalization error) and is easier to compute. Experiments with linear regression and deep neural networks show that DDL also outperforms cross-validation.