Differential Description Length for Hyperparameter Selection in Supervised Learning

Differential Description Length for Hyperparameter Selection in Supervised Learning
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
2020 International Symposium on Information Theory and Its Applications (ISITA)
影响因子:
--
通讯作者:
M. Abolfazli;A. Høst-Madsen;June Zhang
M. Abolfazli;A. Høst-Madsen;June Zhang
中科院分区:
其他
文献类型:
--
作者:
M. Abolfazli;A. Høst-Madsen;June Zhang

文献摘要

相似文献

最小的描述长度(MDL)是用于监督学习问题的模型选择的一种,经常用于模型选择。降低性能的统计假设; MDL的旨在最小化概括的误差。它计算,反映出“旧数据”的“新”数据的条件概率。与MDL相比,要使用整个数据进行验证和测试。也优于交叉验证。
Minimum description length (MDL) is an established method for model selection. For supervised learning problems, cross-validation is often used for model selection in practice. Reasons are 1) MDL is difficult to apply directly to data; 2) MDL may make restrictive statistical assumptions that decrease performance; and 3) MDL does not directly aim to minimize generalization error. In this paper, we introduce a modification to MDL, which we call differential description length (DDL). DDL partitions the data so that the codelength(s) it computes, reflects the conditional probability of seeing ‘new’ data given ‘old’ data. This differential codelength is what allows DDL to estimate generalization error like cross-validation. DDL is also better than cross-validation because it allows the learning algorithm to use the entire data without having to withhold subsets for validation and testing. Compared with MDL, DDL has both better performance (in finding models with smaller generalization error) and is easier to compute. Experiments with linear regression and deep neural networks show that DDL also outperforms cross-validation.