Summarizing Source Code from Structure and Context

Summarizing Source Code from Structure and Context
复制标题

DOI:
10.1109/ijcnn55064.2022.9892013
复制
发表时间:
2022-07
期刊:
2022 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Shifu Hou;Lingwei Chen;Yanfang Ye
Shifu Hou;Lingwei Chen;Yanfang Ye
中科院分区:
其他
文献类型:
--
作者:
Shifu Hou;Lingwei Chen;Yanfang Ye

文献摘要

相似文献

现代软件开发人员倾向于使用社交编码平台来重用代码片段以加快开发过程,而这些平台上的代码通常会受到注释不匹配、丢失或过时的影响。这使得代码搜索和理解变得困难,并且增加了基于这些代码构建的软件的维护负担。由于总结代码是有益的,但它是非常昂贵的手动操作,在本文中,我们阐述了一个自动和有效的代码摘要范例,以解决这一艰巨的挑战。我们将给定的代码片段表示为抽象语法树(AST),并生成一组组合根到叶路径,以使AST可以以不太复杂但富有表现力的方式访问代码上下文和结构。因此,我们设计了一个基于树的Transformer模型,称为TreeXFMR,在这些路径上总结源代码的分层注意力操作。这对代码表示学习产生了两个优点:(1)标记和路径级别的注意机制从不同方面关注源代码的语义和交互;(2)引入的双层位置编码揭示了AST的内部和路径间结构,提高了表示的无歧义性。在解码过程中,TreeXFMR参与这样的学习表示,以产生自然语言单词的每个输出。我们进一步预训练Transformer,以实现更快更好的训练收敛结果。对GitHub上的代码集合进行的大量实验证明了TreeXFMR的有效性,它的性能明显优于最先进的基线。
Modern software developers tend to engage in social coding platforms to reuse code snippets to expedite the development process, while the codes on such platforms are often suffering from comments being mismatched, missing or outdated. This puts the code search and comprehension in difficulty, and increases the burden of maintenance for software building upon these codes. As summarizing code is beneficial yet it is very expensive for manual operation, in this paper, we elaborate an automatic and effective code summarization paradigm to address this laborious challenge. We represent a given code snippet as an abstract syntax tree (AST), and generate a set of compositional root-to-leaf paths to make the AST accessible regarding code context and structure in a less complex yet expressive way. Accordingly, we design a tree-based transformer model, called TreeXFMR, on these paths to summarize source code in a hierarchical attention operation. This yields two advantages on code representation learning: (1) attention mechanisms at token-and path-level attend the semantics and interactions of source code from different aspects; (2) bi-level positional encodings introduced reveal the intra- and inter-path structure of AST and improve the unambiguity of the representations. During decoding, TreeXFMR attends such learned representations to produce each output of natural language word. We further pre-train the transformer to achieve faster and better training convergence results. Extensive experiments on the code collection from GitHub demonstrate the effectiveness of TreeXFMR, which significantly outperforms state-of-the-art baselines.