Impact of Evaluation Methodologies on Code Summarization

Impact of Evaluation Methodologies on Code Summarization
复制标题

DOI:
10.18653/v1/2022.acl-long.339
复制
发表时间:
2021-08
期刊:
--
影响因子:
--
通讯作者:
Pengyu Nie;Jiyang Zhang;Junyi Jessy Li;R. Mooney;Miloš Gligorić
Pengyu Nie;Jiyang Zhang;Junyi Jessy Li;R. Mooney;Miloš Gligorić
中科院分区:
其他
文献类型:
--
作者:
Pengyu Nie;Jiyang Zhang;Junyi Jessy Li;R. Mooney;Miloš Gligorić

文献摘要

相似文献

人们对开发用于代码总结任务的机器学习(ML)模型越来越感兴趣,例如,注释生成和方法命名。尽管机器学习模型的有效性大幅提高,但评估方法,即人们将数据集分成训练集、验证集和测试集的方式,并没有得到很好的研究。具体地说,之前的代码总结工作没有在评估期间考虑代码和注释的时间戳。这可能导致与预期用例不一致的评估。本文介绍了代码摘要研究界的一种新颖的时间分段评估方法,并将其与混合项目和跨项目方法进行了比较。每种方法都可以映射到一些用例,并且在评估ML模型以进行代码总结时应采用时间分段方法。为了评估方法的影响,我们收集了一个带有时间戳的(代码,注释)对数据集,以训练和评估几个用于代码摘要的最新ML模型。我们的实验表明,不同的方法会导致相互矛盾的评估结果。我们邀请社区扩大评估中使用的一套方法。
There has been a growing interest in developing machine learning (ML) models for code summarization tasks, e.g., comment generation and method naming. Despite substantial increase in the effectiveness of ML models, the evaluation methodologies, i.e., the way people split datasets into training, validation, and test sets, were not well studied. Specifically, no prior work on code summarization considered the timestamps of code and comments during evaluation. This may lead to evaluations that are inconsistent with the intended use cases. In this paper, we introduce the time-segmented evaluation methodology, which is novel to the code summarization research community, and compare it with the mixed-project and cross-project methodologies that have been commonly used. Each methodology can be mapped to some use cases, and the time-segmented methodology should be adopted in the evaluation of ML models for code summarization. To assess the impact of methodologies, we collect a dataset of (code, comment) pairs with timestamps to train and evaluate several recent ML models for code summarization. Our experiments show that different methodologies lead to conflicting evaluation results. We invite the community to expand the set of methodologies used in evaluations.