SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code Representations

SPT-Code: Sequence-to-Sequence Pre-Training for Learning Source Code Representations
复制标题

DOI:
10.1145/3510003.3510096
复制
发表时间:
2022-01
期刊:
2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Changan Niu;Chuanyi Li;Vincent Ng;Jidong Ge;LiGuo Huang;B. Luo
Changan Niu;Chuanyi Li;Vincent Ng;Jidong Ge;LiGuo Huang;B. Luo
中科院分区:
其他
文献类型:
--
作者:
Changan Niu;Chuanyi Li;Vincent Ng;Jidong Ge;LiGuo Huang;B. Luo

文献摘要

相似文献

近年来,大型预训练模型成功应用于代码表示学习,从而使许多与代码相关的下游任务得到了实质性改进。但它们在 SE 任务中的应用存在一些问题。首先,大多数预训练模型只专注于预训练 Transformer 的编码器。然而,对于使用具有编码器-解码器架构的模型解决的生成任务,没有理由在预训练期间忽略解码器。其次,许多现有的预训练模型,包括 T5 学习等最先进的模型,只是重复使用为自然语言设计的预训练任务。此外,为了学习代码摘要等代码相关任务最终所需的源代码的自然语言描述,现有的预训练任务需要由源代码和相关自然语言描述组成的双语语料库,这严重限制了预训练的数据量。为此,我们提出了 SPT-Code,一种针对源代码的序列到序列预训练模型。为了以序列到序列的方式预训练SPT-Code,并解决现有预训练任务的上述弱点,我们引入了三个预训练任务,专门设计用于使SPT-Code能够在不依赖任何双语语料库的情况下学习源代码知识、相应的代码结构以及代码的自然语言描述,并最终在应用于下游任务时利用这三个信息源。实验结果表明,经过微调,SPT-Code 在五个与代码相关的下游任务上实现了最先进的性能。
Recent years have seen the successful application of large pretrained models to code representation learning, resulting in substantial improvements on many code-related downstream tasks. But there are issues surrounding their application to SE tasks. First, the majority of the pre-trained models focus on pre-training only the encoder of the Transformer. For generation tasks that are addressed using models with the encoder-decoder architecture, however, there is no reason why the decoder should be left out during pre-training. Second, many existing pre-trained models, including state-of-the-art models such as T5-learning, simply reuse the pretraining tasks designed for natural languages. Moreover, to learn the natural language description of source code needed eventually for code-related tasks such as code summarization, existing pretraining tasks require a bilingual corpus composed of source code and the associated natural language description, which severely limits the amount of data for pre-training. To this end, we propose SPT-Code, a sequence-to-sequence pre-trained model for source code. In order to pre-train SPT-Code in a sequence-to-sequence manner and address the aforementioned weaknesses associated with existing pre-training tasks, we introduce three pre-training tasks that are specifically designed to enable SPT-Code to learn knowledge of source code, the corresponding code structure, as well as a natural language description of the code without relying on any bilingual corpus, and eventually exploit these three sources of information when it is applied to downstream tasks. Experimental results demonstrate that SPT-Code achieves state-of-the-art performance on five code-related downstream tasks after fine-tuning.