Code to Comment Translation: A Comparative Study on Model Effectiveness & Errors

Code to Comment Translation: A Comparative Study on Model Effectiveness & Errors
复制标题

DOI:
10.18653/v1/2021.nlp4prog-1.1
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Junayed Mahmud;FAHIM FAISAL;Raihan Islam Arnob;Antonios Anastasopoulos;Kevin Moran
Junayed Mahmud;FAHIM FAISAL;Raihan Islam Arnob;Antonios Anastasopoulos;Kevin Moran
中科院分区:
其他
文献类型:
--
作者:
Junayed Mahmud;FAHIM FAISAL;Raihan Islam Arnob;Antonios Anastasopoulos;Kevin Moran

文献摘要

相似文献

自动代码摘要是一个流行的软件工程研究主题,其中使用机器翻译模型将代码片段“转换为相关的自然语言描述,则使用基于自动参考的指标进行了大多数评估。我们认为,编程语言和自然语言之间的差距,这一研究线将受益于对各种的定性投资因此,当前最新模型的误差模式。 -4,流星和胭脂l机器翻译指标,在我们的定性评估中,我们对模型最常见的错误进行了手动开放编码,与地面真相相比字幕。我们的投资揭示了基于度量的绩效与模型预测错误之间的新见解,该错误基于错误分类法,可用于推动未来的研究工作。
Automated source code summarization is a popular software engineering research topic wherein machine translation models are employed to “translate” code snippets into relevant natural language descriptions. Most evaluations of such models are conducted using automatic reference-based metrics. However, given the relatively large semantic gap between programming languages and natural language, we argue that this line of research would benefit from a qualitative investigation into the various error modes of current state-of-the-art models. Therefore, in this work, we perform both a quantitative and qualitative comparison of three recently proposed source code summarization models. In our quantitative evaluation, we compare the models based on the smoothed BLEU-4, METEOR, and ROUGE-L machine translation metrics, and in our qualitative evaluation, we perform a manual open-coding of the most common errors committed by the models when compared to ground truth captions. Our investigation reveals new insights into the relationship between metric-based performance and model prediction errors grounded in an error taxonomy that can be used to drive future research efforts.