Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks

Evaluation of Neural Architectures Trained with Square Loss vs Cross-Entropy in Classification Tasks
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Like Hui;M. Belkin
Like Hui;M. Belkin
中科院分区:
其他
文献类型:
--
作者:
Like Hui;M. Belkin

文献摘要

相似文献

用于分类任务的现代神经架构是使用交叉熵损失来训练的,交叉熵损失被广泛认为在经验上上级平方损失。在这项工作中,我们提供的证据表明,这种信念可能是没有根据的。我们探索了几种主要的神经架构和一系列用于NLP,自动语音识别(ASR)和计算机视觉任务的标准基准数据集,以表明这些架构具有与文献中报道的相同的超参数设置,即使在均衡计算资源之后,在平方损失训练时也表现出色或更好。事实上,我们观察到平方损失在绝大多数NLP和ASR实验中产生更好的结果。交叉熵似乎在计算机视觉任务上有轻微的优势。我们认为,几乎没有令人信服的经验或理论证据表明一个明确的优势,交叉熵损失。事实上,在我们的实验中,几乎所有非视觉任务的表现都可以通过切换到平方损失来提高,有时甚至是显着提高。此外,平方损失训练似乎对初始化的随机性不太敏感。我们认为,使用平方损失进行分类的训练需要成为现代深度学习最佳实践的一部分,与交叉熵处于同等地位。
Modern neural architectures for classification tasks are trained using the cross-entropy loss, which is widely believed to be empirically superior to the square loss. In this work we provide evidence indicating that this belief may not be well-founded. We explore several major neural architectures and a range of standard benchmark datasets for NLP, automatic speech recognition (ASR) and computer vision tasks to show that these architectures, with the same hyper-parameter settings as reported in the literature, perform comparably or better when trained with the square loss, even after equalizing computational resources. Indeed, we observe that the square loss produces better results in the dominant majority of NLP and ASR experiments. Cross-entropy appears to have a slight edge on computer vision tasks. We argue that there is little compelling empirical or theoretical evidence indicating a clear-cut advantage to the cross-entropy loss. Indeed, in our experiments, performance on nearly all non-vision tasks can be improved, sometimes significantly, by switching to the square loss. Furthermore, training with square loss appears to be less sensitive to the randomness in initialization. We posit that training using the square loss for classification needs to be a part of best practices of modern deep learning on equal footing with cross-entropy.