Label Smoothing Improves Neural Source Code Summarization

Label Smoothing Improves Neural Source Code Summarization
复制标题

DOI:
10.1109/icpc58990.2023.00025
复制
发表时间:
2023-03
期刊:
2023 IEEE/ACM 31st International Conference on Program Comprehension (ICPC)
影响因子:
--
通讯作者:
S. Haque;Aakash Bansal;Collin McMillan
S. Haque;Aakash Bansal;Collin McMillan
中科院分区:
其他
文献类型:
--
作者:
S. Haque;Aakash Bansal;Collin McMillan

文献摘要

相似文献

标签平滑是神经网络的一种正则化技术。通常,神经模型被训练成一个输出分布,该分布是一个向量,其中1表示正确的预测,0表示所有其他元素。标签平滑将正确的预测位置转换为略小于1的值,然后将余数分配给其他元素,使它们略大于0。标签平滑背后的一个概念性解释是,它有助于防止神经模型变得“过度自信”,迫使它考虑替代方案,即使只是轻微的。标签平滑已被证明有助于语言生成的几个领域,但通常需要大量的调优和测试才能实现最佳结果。这种调优和测试还没有被用于神经源代码摘要的报道,神经源代码摘要是软件工程中一个不断发展的研究领域,旨在生成源代码行为的自然语言描述。在本文中,我们展示了标签平滑对神经代码摘要中几个基线的影响,并进行了一项实验,以找到标签平滑的好参数,并对其使用提出建议。
Label smoothing is a regularization technique for neural networks. Normally neural models are trained to an output distribution that is a vector with a single 1 for the correct prediction, and 0 for all other elements. Label smoothing converts the correct prediction location to something slightly less than 1, then distributes the remainder to the other elements such that they are slightly greater than 0. A conceptual explanation behind label smoothing is that it helps prevent a neural model from becoming "overconfident" by forcing it to consider alternatives, even if only slightly. Label smoothing has been shown to help several areas of language generation, yet typically requires considerable tuning and testing to achieve the optimal results. This tuning and testing has not been reported for neural source code summarization – a growing research area in software engineering that seeks to generate natural language descriptions of source code behavior. In this paper, we demonstrate the effect of label smoothing on several baselines in neural code summarization, and conduct an experiment to find good parameters for label smoothing and make recommendations for its use.