Early-stopped neural networks are consistent

Early-stopped neural networks are consistent
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Ziwei Ji;Justin D. Li;Matus Telgarsky
Ziwei Ji;Justin D. Li;Matus Telgarsky
中科院分区:
其他
文献类型:
--
作者:
Ziwei Ji;Justin D. Li;Matus Telgarsky

文献摘要

被引文献

相似文献

这项工作研究了通过梯度下降对二进制分类数据进行逻辑损失训练的浅层ReLU网络的行为,其中底层数据分布是一般的,并且(最佳)贝叶斯风险不一定为零。在这种情况下,它表明,早期停止的梯度下降实现人口风险任意接近最优,不仅在逻辑和误分类损失方面,而且在校准方面,这意味着其输出的S形映射任意精细地近似真实的潜在条件分布。此外,这种分析所必需的迭代、样本和架构复杂性都可以自然地随着真实条件模型的某种复杂性度量而扩展。最后,虽然它没有表明,早期停止是必要的,它表明,任何一个单变量分类满足本地插值属性是不一致的。
This work studies the behavior of shallow ReLU networks trained with the logistic loss via gradient descent on binary classification data where the underlying data distribution is general, and the (optimal) Bayes risk is not necessarily zero. In this setting, it is shown that gradient descent with early stopping achieves population risk arbitrarily close to optimal in terms of not just logistic and misclassification losses, but also in terms of calibration, meaning the sigmoid mapping of its outputs approximates the true underlying conditional distribution arbitrarily finely. Moreover, the necessary iteration, sample, and architectural complexities of this analysis all scale naturally with a certain complexity measure of the true conditional model. Lastly, while it is not shown that early stopping is necessary, it is shown that any univariate classifier satisfying a local interpolation property is inconsistent.