Contrastive Code Representation Learning

Contrastive Code Representation Learning
复制标题

DOI:
10.18653/v1/2021.emnlp-main.482
复制
发表时间:
2020-07
期刊:
--
影响因子:
--
通讯作者:
Paras Jain;Ajay Jain;Tianjun Zhang;P. Abbeel;Joseph Gonzalez;I. Stoica
Paras Jain;Ajay Jain;Tianjun Zhang;P. Abbeel;Joseph Gonzalez;I. Stoica
中科院分区:
其他
文献类型:
--
作者:
Paras Jain;Ajay Jain;Tianjun Zhang;P. Abbeel;Joseph Gonzalez;I. Stoica

文献摘要

被引文献

相似文献

最近的工作通过从其上下文中重建代币来了解源代码的上下文表示。对于下游语义理解诸如代码克隆检测之类的任务,这些表示形式应理想地捕获程序功能。但是,我们表明,即使在编辑保留语义时,流行的基于重建的Roberta模型也对源代码编辑也很敏感。我们提出了违反:学习代码功能而不是形式的对比前训练任务。在许多非等效的干扰因素中,降低了培训的神经网络,以识别程序的功能相似变体。我们使用自动化的源对源编译器作为数据增强形式可靠地生成这些变体。对比预训练在对抗代码克隆检测基准测试基准中优于罗伯塔(Roberta)的表现为39%的AUROC。令人惊讶的是,改善的对抗性鲁棒性转化为比自然代码的精度更好。在竞争基线的基线上,Contracode将摘要和打字条类型推理精度提高了2至13个百分点。所有来源均可在https://github.com/parasj/contracode上找到。
Recent work learns contextual representations of source code by reconstructing tokens from their context. For downstream semantic understanding tasks like code clone detection, these representations should ideally capture program functionality. However, we show that the popular reconstruction-based RoBERTa model is sensitive to source code edits, even when the edits preserve semantics. We propose ContraCode: a contrastive pre-training task that learns code functionality, not form. ContraCode pre-trains a neural network to identify functionally similar variants of a program among many non-equivalent distractors. We scalably generate these variants using an automated source-to-source compiler as a form of data augmentation. Contrastive pre-training outperforms RoBERTa on an adversarial code clone detection benchmark by 39% AUROC. Surprisingly, improved adversarial robustness translates to better accuracy over natural code; ContraCode improves summarization and TypeScript type inference accuracy by 2 to 13 percentage points over competitive baselines. All source is available at https://github.com/parasj/contracode.