STATISTICAL-THEORY OF LEARNING-CURVES UNDER ENTROPIC LOSS CRITERION

STATISTICAL-THEORY OF LEARNING-CURVES UNDER ENTROPIC LOSS CRITERION
复制标题

DOI:
10.1162/neco.1993.5.1.140
复制
发表时间:
1993-01-01
期刊:
影响因子:
2.9
通讯作者:
MURATA, N
MURATA, N
中科院分区:
计算机科学4区
文献类型:
--
作者:
AMARI, S;MURATA, N

文献摘要

被引文献

相似文献

本文阐明了学习曲线的一个普遍性质,它显示了泛化误差,训练误差和底层随机机器的复杂性是如何相关的,以及随机机器的行为如何随着训练示例数量的增加而改善。误差由熵损失来度量。证明了泛化误差收敛于真实机器条件分布的熵H 0,即H 0 + m*/(2 t),而训练误差收敛于H 0- m*/(2 t),其中t为样本数,m* 表示网络的复杂度。当模型是忠实的,意味着真正的机器在模型中,m* 减少到m,可修改的参数的数量。这是一个普遍的规律,因为它适用于任何规则机器,而不管它在最大似然估计下的结构如何。贝叶斯和吉布斯学习算法得到类似的关系。这些学习曲线显示了学习的准确性,模型的复杂性和训练样本数量之间的关系。
The present paper elucidates a universal property of learning curves, which shows how the generalization error, training error, and the complexity of the underlying stochastic machine are related and how the behavior of a stochastic machine is improved as the number of training examples increases. The error is measured by the entropic loss. It is proved that the generalization error converges to H0, the entropy of the conditional distribution of the true machine, as H0 + m*/(2t), while the training error converges as H0 - m*/(2t), where t is the number of examples and m* shows the complexity of the network. When the model is faithful, implying that the true machine is in the model, m* is reduced to m, the number of modifiable parameters. This is a universal law because it holds for any regular machine irrespective of its structure under the maximum likelihood estimator. Similar relations are obtained for the Bayes and Gibbs learning algorithms. These learning curves show the relation among the accuracy of learning, the complexity of a model, and the number of training examples.