The minimum description length principle in coding and modeling

The minimum description length principle in coding and modeling
复制标题

DOI:
10.1109/18.720554
复制
发表时间:
1998-10-01
影响因子:
2.5
通讯作者:
Yu, B
Yu, B
中科院分区:
计算机科学2区
文献类型:
--
作者:
Barron, A;Rissanen, J;Yu, B

文献摘要

被引文献

相似文献

我们回顾了最小描述长度和随机复杂性的原则,用于数据压缩和统计建模。随机复杂度被制定为最佳通用编码问题的解决方案,扩展香农的基本信源编码定理。归一化的最大似然,混合物,和预测编码的每个示出实现随机复杂性内渐近消失的条款。我们从数据压缩质量和统计推断准确性的Vantage来评估最小描述长度标准的性能。上下文树建模、密度估计和高斯线性回归中的模型选择可作为示例。
We review the principles of Minimum Description Length and Stochastic Complexity as used in data compression and statistical modeling. Stochastic complexity is formulated as the solution to optimum universal coding problems extending Shannon's basic source coding theorem. The normalized maximized likelihood, mixture, and predictive codings are each shown to achieve the stochastic complexity to within asymptotically vanishing terms. We assess the performance of the minimum description length criterion both from the vantage point of quality of data compression and accuracy of statistical inference. Context tree modeling, density estimation, and model selection in Gaussian linear regression serve as examples.