Semantic Similarity Metrics for Evaluating Source Code Summarization

Semantic Similarity Metrics for Evaluating Source Code Summarization
复制标题

DOI:
10.1145/nnnnnnn.nnnnnnn
复制
发表时间:
2017-05
期刊:
2022 IEEE/ACM 30th International Conference on Program Comprehension (ICPC)
影响因子:
--
通讯作者:
Yang Liu;Goran Radanovic;Christos Dimitrakakis;Debmalya Mandal;D. Parkes
Yang Liu;Goran Radanovic;Christos Dimitrakakis;Debmalya Mandal;D. Parkes
中科院分区:
其他
文献类型:
--
作者:
Yang Liu;Goran Radanovic;Christos Dimitrakakis;Debmalya Mandal;D. Parkes

文献摘要

被引文献

相似文献

源代码摘要涉及用自然语言创建源代码的简要描述。这些描述是软件文档的关键组成部分,例如JavaBean。自动代码摘要是软件工程研究的一个重要目标,因为摘要对程序员具有很高的价值,同时手工编写和维护文档的成本也很高。目前的工作几乎都是基于通过大数据输入训练的机器模型。代码的示例和该代码的摘要的大数据集用于训练例如编码器-解码器神经模型。然后,根据一组参考摘要评估模型的输出预测。输入是模型看不到的代码,预测与参考进行比较。将预测与参考进行比较的方法基本上是通过BLEU或ROUGE等度量计算的单词重叠。使用单词重叠的问题是,不是一个句子中的所有单词都具有相同的重要性,并且许多单词都有同义词。结果是计算的相似性可能与人类读者感知的相似性不匹配。在本文中,我们进行了一项实验,以衡量各种词重叠指标与预测和参考摘要的人类评级相似性的程度。我们评估替代品的语义相似性度量的基础上,目前的工作,并提出建议,评估源代码摘要。
Source code summarization involves creating brief descriptions of source code in natural language. These descriptions are a key component of software documentation such as JavaDocs. Automatic code summarization is a prized target of software engineering research, due to the high value summaries have to programmers and the simultaneously high cost of writing and maintaining documentation by hand. Current work is almost all based on machine models trained via big data input. Large datasets of examples of code and summaries of that code are used to train an e.g. encoder-decoder neural model. Then the output predictions of the model are evaluated against a set of reference summaries. The input is code not seen by the model, and the prediction is compared to a reference. The means by which a prediction is compared to a reference is essentially word overlap, calculated via a metric such as BLEU or ROUGE. The problem with using word overlap is that not all words in a sentence have the same importance, and many words have synonyms. The result is that calculated similarity may not match the perceived similarity by human readers. In this paper, we conduct an experiment to measure the degree to which various word overlap metrics correlate to human-rated similarity of predicted and reference summaries. We evaluate alternatives based on current work in semantic similarity metrics and propose recommendations for evaluation of source code summarization.