Towards Learning (Dis)-Similarity of Source Code from Program Contrasts

Towards Learning (Dis)-Similarity of Source Code from Program Contrasts
复制标题

DOI:
10.18653/v1/2022.acl-long.436
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Yangruibo Ding;Luca Buratti;Saurabh Pujar;Alessandro Morari;Baishakhi Ray;Saikat Chakraborty
Yangruibo Ding;Luca Buratti;Saurabh Pujar;Alessandro Morari;Baishakhi Ray;Saikat Chakraborty
中科院分区:
其他
文献类型:
--
作者:
Yangruibo Ding;Luca Buratti;Saurabh Pujar;Alessandro Morari;Baishakhi Ray;Saikat Chakraborty

文献摘要

相似文献

理解源代码的功能相似性对于软件漏洞和代码克隆检测等代码建模任务具有重要意义。我们提出了DISCO(DIS相似性的代码),一种新的自我监督模型,专注于识别(DIS)类似的功能的源代码。与现有的工作不同,我们的方法不需要大量的随机收集的数据集。相反,我们设计了结构引导的代码转换算法来生成合成代码克隆并注入真实世界的安全漏洞,以有针对性的方式增强收集的数据集。我们建议使用这种自动生成的程序对比来预训练Transformer模型,以更好地识别野生环境中的类似代码,并将易受攻击的程序与良性程序区分开来。为了更好地捕捉源代码的结构特征,我们提出了一个新的完形填空目标来编码基于局部树的上下文(例如,父母或兄弟节点)。我们用一个小得多的数据集来预训练我们的模型,其大小仅为最先进模型训练数据集的5%,以说明我们的数据增强和预训练方法的有效性。评估表明,即使数据少得多,DISCO仍然可以在漏洞和代码克隆检测任务中超越最先进的模型。
Understanding the functional (dis)-similarity of source code is significant for code modeling tasks such as software vulnerability and code clone detection. We present DISCO (DIS-similarity of COde), a novel self-supervised model focusing on identifying (dis)similar functionalities of source code. Different from existing works, our approach does not require a huge amount of randomly collected datasets. Rather, we design structure-guided code transformation algorithms to generate synthetic code clones and inject real-world security bugs, augmenting the collected datasets in a targeted way. We propose to pre-train the Transformer model with such automatically generated program contrasts to better identify similar code in the wild and differentiate vulnerable programs from benign ones. To better capture the structural features of source code, we propose a new cloze objective to encode the local tree-based context (e.g., parents or sibling nodes). We pre-train our model with a much smaller dataset, the size of which is only 5% of the state-of-the-art models’ training datasets, to illustrate the effectiveness of our data augmentation and the pre-training approach. The evaluation shows that, even with much less data, DISCO can still outperform the state-of-the-art models in vulnerability and code clone detection tasks.