Function Call Graph Context Encoding for Neural Source Code Summarization

Function Call Graph Context Encoding for Neural Source Code Summarization
复制标题

DOI:
10.1109/tse.2023.3279774
复制
发表时间:
2023-09
影响因子:
7.4
通讯作者:
Aakash Bansal;Zachary Eberhart;Z. Karas;Yu Huang;Collin McMillan
Aakash Bansal;Zachary Eberhart;Z. Karas;Yu Huang;Collin McMillan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Aakash Bansal;Zachary Eberhart;Z. Karas;Yu Huang;Collin McMillan

文献摘要

相似文献

源代码摘要是编写源代码的自然语言描述的任务。这些描述的主要用途是在程序员的文档中。由于程序员自己编写这些描述的时间成本,这些描述的自动生成是一个高价值的研究目标。近年来,软件工程和人工智能研究的融合,通过应用源代码的神经模型,已经进入了源代码自动摘要领域。然而,绝大多数方法的一个致命弱点是,它们往往只依赖于被总结的源代码所提供的上下文。但程序理解中的经验研究非常清楚,描述代码所需的信息更多地驻留在代码周围的函数调用图形式的上下文中。在本文中,我们提出了一种编码这种调用图上下文的技术,用于代码摘要的神经模型。我们将我们的方法作为现有方法的补充来实现,并显示出与现有方法相比在统计上的显著改进。在一项对20名程序员进行的人类研究中,我们表明,程序员认为生成的摘要通常与人类编写的摘要一样准确、可读和简洁。
Source code summarization is the task of writing natural language descriptions of source code. The primary use of these descriptions is in documentation for programmers. Automatic generation of these descriptions is a high value research target due to the time cost to programmers of writing these descriptions themselves. In recent years, a confluence of software engineering and artificial intelligence research has made inroads into automatic source code summarization through applications of neural models of that source code. However, an Achilles’ heel to a vast majority of approaches is that they tend to rely solely on the context provided by the source code being summarized. But empirical studies in program comprehension are quite clear that the information needed to describe code much more often resides in the context in the form of Function Call Graph surrounding that code. In this paper, we present a technique for encoding this call graph context for neural models of code summarization. We implement our approach as a supplement to existing approaches, and show statistically significant improvement over existing approaches. In a human study with 20 programmers, we show that programmers perceive generated summaries to generally be as accurate, readable, and concise as human-written summaries.