Model Similarity Mitigates Test Set Overuse

Model Similarity Mitigates Test Set Overuse
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Horia Mania;John Miller;Ludwig Schmidt;Moritz Hardt;B. Recht
Horia Mania;John Miller;Ludwig Schmidt;Moritz Hardt;B. Recht
中科院分区:
其他
文献类型:
--
作者:
Horia Mania;John Miller;Ludwig Schmidt;Moritz Hardt;B. Recht

文献摘要

被引文献

相似文献

在当今的机器学习工作流中,过度重用测试数据已经变得司空见惯。流行的基准测试、竞赛、工业规模调整,以及其他应用程序,都涉及超出统计置信限的测试数据重用。尽管如此,最近的复制研究表明,尽管经过多年的广泛重复使用,但流行的基准继续支持进步。我们为测试数据的表面寿命提供了一种新的解释:许多提出的模型在预测上是相似的,我们证明这种相似性缓解了过度拟合。具体地说,我们的经验表明,为ImageNet ILSVRC基准提出的模型与其预测的一致性远远超出了我们仅从其精度水平所能得出的结论。同样,通过大规模超参数搜索创建的模型具有很高的相似性。在这些经验观察的启发下,我们给出了一个考虑了相似性的非渐近推广界,从而在实际环境中得到了有意义的置信界。
Excessive reuse of test data has become commonplace in today's machine learning workflows. Popular benchmarks, competitions, industrial scale tuning, among other applications, all involve test data reuse beyond guidance by statistical confidence bounds. Nonetheless, recent replication studies give evidence that popular benchmarks continue to support progress despite years of extensive reuse. We proffer a new explanation for the apparent longevity of test data: Many proposed models are similar in their predictions and we prove that this similarity mitigates overfitting. Specifically, we show empirically that models proposed for the ImageNet ILSVRC benchmark agree in their predictions well beyond what we can conclude from their accuracy levels alone. Likewise, models created by large scale hyperparameter search enjoy high levels of similarity. Motivated by these empirical observations, we give a non-asymptotic generalization bound that takes similarity into account, leading to meaningful confidence bounds in practical settings.