RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style Transformation

RoPGen: Towards Robust Code Authorship Attribution via Automatic Coding Style Transformation
复制标题

DOI:
10.1145/3510003.3510181
复制
发表时间:
2022-02
期刊:
2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Zhen Li;Guenevere Chen; Chen-Chen-Chen;Yayi Zou;Shouhuai Xu
Zhen Li;Guenevere Chen; Chen-Chen-Chen;Yayi Zou;Shouhuai Xu
中科院分区:
其他
文献类型:
--
作者:
Zhen Li;Guenevere Chen; Chen-Chen-Chen;Yayi Zou;Shouhuai Xu

文献摘要

被引文献

相似文献

源代码作者归属是软件取证、缺陷修复和软件质量分析等应用中经常遇到的重要问题。最近的研究表明,当前的源代码作者归属方法可能会受到攻击者利用对抗性示例和编码风格操纵的影响。这就需要对代码作者归属问题的鲁棒解决方案。在本文中,我们开始研究基于深度学习(DL)的代码作者归属鲁棒性。我们提出了一个创新的框架称为鲁棒编码风格模式生成(RoPGen),它本质上是学习作者的独特的编码风格模式,攻击者很难操纵或模仿。其关键思想是在对抗训练阶段将联合收割机数据增强和梯度增强相结合。这有效地增加了训练示例的多样性,对深度神经网络的梯度产生了有意义的扰动,并学习了编码风格的多样化表示。我们使用四个数据集的C,C++和Java编写的程序的有效性进行评估的RoPGen。实验结果表明,RoPGen能显著提高基于DL的代码作者归属的鲁棒性,分别降低了22.8%和41.0%的有针对性攻击和无针对性攻击的成功率.
Source code authorship attribution is an important problem often encountered in applications such as software forensics, bug fixing, and software quality analysis. Recent studies show that current source code authorship attribution methods can be compromised by attackers exploiting adversarial examples and coding style ma-nipulation. This calls for robust solutions to the problem of code authorship attribution. In this paper, we initiate the study on making Deep Learning (DL)-based code authorship attribution robust. We propose an innovative framework called Robust coding style Patterns Generation (RoPGen), which essentially learns authors' unique coding style patterns that are hard for attackers to manip-ulate or imitate. The key idea is to combine data augmentation and gradient augmentation at the adversarial training phase. This effectively increases the diversity of training examples, generates meaningful perturbations to gradients of deep neural networks, and learns diversified representations of coding styles. We evaluate the effectiveness of RoPGen using four datasets of programs written in C, C++, and Java. Experimental results show that RoPGen can significantly improve the robustness of DL-based code authorship attribution, by respectively reducing 22.8% and 41.0% of the success rate of targeted and untargeted attacks on average.