Language-Agnostic Representation Learning of Source Code from Structure and Context

Language-Agnostic Representation Learning of Source Code from Structure and Context
复制标题

DOI:
--
复制
发表时间:
2021-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Daniel Zugner;Tobias Kirschstein;Michele Catasta;J. Leskovec;Stephan Gunnemann
Daniel Zugner;Tobias Kirschstein;Michele Catasta;J. Leskovec;Stephan Gunnemann
中科院分区:
其他
文献类型:
--
作者:
Daniel Zugner;Tobias Kirschstein;Michele Catasta;J. Leskovec;Stephan Gunnemann

文献摘要

被引文献

相似文献

源代码(上下文)及其解析的抽象语法树(AST;结构)是同一计算机程序的两个互补表示。传统上,机器学习模型的设计者主要依赖于结构或上下文。我们提出了一个新的模型,它联合学习上下文和源代码的结构。与以前的方法相比,我们的模型只使用语言无关的功能,即,可以直接从AST计算的源代码和功能。除了获得国家的最先进的单语代码摘要在这项工作中考虑的所有五种编程语言,我们提出了第一个多语言代码摘要模型。我们发现,对来自多种编程语言的非并行数据进行联合训练可以改善所有语言的结果,其中最大的收益是低资源语言。值得注意的是,仅从上下文进行多语言训练并没有带来同样的改进,突出了将结构和上下文结合起来进行代码表示学习的好处。
Source code (Context) and its parsed abstract syntax tree (AST; Structure) are two complementary representations of the same computer program. Traditionally, designers of machine learning models have relied predominantly either on Structure or Context. We propose a new model, which jointly learns on Context and Structure of source code. In contrast to previous approaches, our model uses only language-agnostic features, i.e., source code and features that can be computed directly from the AST. Besides obtaining state-of-the-art on monolingual code summarization on all five programming languages considered in this work, we propose the first multilingual code summarization model. We show that jointly training on non-parallel data from multiple programming languages improves results on all individual languages, where the strongest gains are on low-resource languages. Remarkably, multilingual training only from Context does not lead to the same improvements, highlighting the benefits of combining Structure and Context for representation learning on code.