Deep Learning to Detect Redundant Method Comments

Deep Learning to Detect Redundant Method Comments
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Annie Louis;Santanu Kumar Dash;Earl T. Barr;Charles Sutton
Annie Louis;Santanu Kumar Dash;Earl T. Barr;Charles Sutton
中科院分区:
其他
文献类型:
--
作者:
Annie Louis;Santanu Kumar Dash;Earl T. Barr;Charles Sutton

文献摘要

被引文献

相似文献

软件中的注释对于维护和重用至关重要。但除了规范性的建议,几乎没有实际的支持或量化的理解是什么使评论有用。在本文中,我们介绍了识别注释的任务,这些注释对它们要记录的代码没有信息。为了解决这个问题,我们从代码中引入了注释蕴涵的概念,高蕴涵表示注释的自然语言语义可以直接从代码中推断出来。虽然不是所有的隐含注释都是低质量的,但是太容易推断的注释,例如,重述代码的注释,被软件风格的权威人士广泛禁止。在此基础上,我们开发了一个名为CRAIC的工具,评分方法级的冗余评论。然后,开发人员可以扩展或删除高度冗余的注释。CRAIC使用深层语言模型来利用大型软件语料库,而无需昂贵的蕴涵手动注释。我们表明,CRAIC可以执行的评论蕴涵任务与人类的判断很好的协议。我们的研究结果也对文档工具产生了影响。例如,我们发现Javadoc中的常见标记从代码中的可预测性至少是非Javadoc语句的两倍,这表明Javadoc标记的信息量比更自由的注释少
Comments in software are critical for maintenance and reuse. But apart from prescriptive advice, there is little practical support or quantitative understanding of what makes a comment useful. In this paper, we introduce the task of identifying comments which are uninformative about the code they are meant to document. To address this problem, we introduce the notion of comment entailment from code, high entailment indicating that a comment's natural language semantics can be inferred directly from the code. Although not all entailed comments are low quality, comments that are too easily inferred, for example, comments that restate the code, are widely discouraged by authorities on software style. Based on this, we develop a tool called CRAIC which scores method-level comments for redundancy. Highly redundant comments can then be expanded or alternately removed by the developer. CRAIC uses deep language models to exploit large software corpora without requiring expensive manual annotations of entailment. We show that CRAIC can perform the comment entailment task with good agreement with human judgements. Our findings also have implications for documentation tools. For example, we find that common tags in Javadoc are at least two times more predictable from code than non-Javadoc sentences, suggesting that Javadoc tags are less informative than more free-form comments