Do We Need Zero Training Loss After Achieving Zero Training Error?

Do We Need Zero Training Loss After Achieving Zero Training Error?
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Takashi Ishida;Ikko Yamane;Tomoya Sakai;Gang Niu;Masashi Sugiyama
Takashi Ishida;Ikko Yamane;Tomoya Sakai;Gang Niu;Masashi Sugiyama
中科院分区:
其他
文献类型:
--
作者:
Takashi Ishida;Ikko Yamane;Tomoya Sakai;Gang Niu;Masashi Sugiyama

文献摘要

被引文献

相似文献

超参数深度网络具有记忆训练数据的能力,训练误差为零。即使在记忆之后,训练损失继续接近于零,使得模型过于自信,测试性能下降。由于现有的正则化子并不直接以避免零训练损失为目标,它们往往不能保持适度的训练损失,最终导致太小或太大的损失。我们提出了一种称为泛洪的直接解决方案,当训练损失达到一个相当小的值时,该解决方案有意地防止进一步减少训练损失,我们称之为泛洪级别。我们的方法通过像往常一样进行小批量梯度下降,使损失在洪泛水位附近浮动,但如果训练损失低于洪泛水位,则进行梯度上升。这可以用一行代码实现,并且与任何随机优化器和其他正则化程序兼容。在洪泛的情况下,模型将继续在相同的非零训练损失下“随机行走”,我们预计它将漂移到具有平坦损失景观的区域,从而导致更好的泛化。我们的实验表明,泛洪提高了性能,并且作为副产品,导致了测试损失的双下降曲线。
Overparameterized deep networks have the capacity to memorize training data with zero training error. Even after memorization, the training loss continues to approach zero, making the model overconfident and the test performance degraded. Since existing regularizers do not directly aim to avoid zero training loss, they often fail to maintain a moderate level of training loss, ending up with a too small or too large loss. We propose a direct solution called flooding that intentionally prevents further reduction of the training loss when it reaches a reasonably small value, which we call the flooding level. Our approach makes the loss float around the flooding level by doing mini-batched gradient descent as usual but gradient ascent if the training loss is below the flooding level. This can be implemented with one line of code, and is compatible with any stochastic optimizer and other regularizers. With flooding, the model will continue to "random walk" with the same non-zero training loss, and we expect it to drift into an area with a flat loss landscape that leads to better generalization. We experimentally show that flooding improves performance and as a byproduct, induces a double descent curve of the test loss.