Semantic Robustness of Models of Source Code

Semantic Robustness of Models of Source Code
复制标题

DOI:
10.1109/saner53432.2022.00070
复制
发表时间:
2020-02
期刊:
2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)
影响因子:
--
通讯作者:
Goutham Ramakrishnan;Jordan Henkel;Zi Wang;Aws Albarghouthi;S. Jha;T. Reps
Goutham Ramakrishnan;Jordan Henkel;Zi Wang;Aws Albarghouthi;S. Jha;T. Reps
中科院分区:
其他
文献类型:
--
作者:
Goutham Ramakrishnan;Jordan Henkel;Zi Wang;Aws Albarghouthi;S. Jha;T. Reps

文献摘要

被引文献

相似文献

深度神经网络容易受到对抗性示例的攻击,可导致预测不正确的扰动。我们为源代码模型研究了此问题,我们希望神经网络对保留代码功能的源代码修改具有鲁棒性。为了促进训练强大的模型,我们定义了一个强大而通用的对手,该对手可以采用参数,语义保护程序转换的序列。然后,我们探索如何借助这样的对手,可以训练对对抗性程序转换的强大模型。我们对我们的方法进行了彻底的评估,并发现了一些令人惊讶的事实:在我们执行的每项评估中,我们发现了强大的培训来击败数据集扩展;我们发现,用于代码模型的最先进的体系结构(Code2Seq)比简单的基线更难使其稳健。此外,我们发现Code2Seq在我们简单的基线模型中没有令人惊讶的弱点。最后,我们发现强大的模型对来自不同来源的看不见的数据的表现更好(正如人们所希望的) - 无论如何,我们还发现,在跨语言转移任务中,健壮的模型并不明显更好。据我们所知,我们是第一个研究代码模型鲁棒性与域适应性和跨语言转移任务之间的相互作用的人。
Deep neural networks are vulnerable to adversarial examples-small input perturbations that result in incorrect predictions. We study this problem for models of source code, where we want the neural network to be robust to source-code modifications that preserve code functionality. To facilitate training robust models, we define a powerful and generic adversary that can employ sequences of parametric, semantics-preserving program transformations. We then explore how, with such an adversary, one can train models that are robust to adversarial program transformations. We conduct a thorough evaluation of our approach and find several surprising facts: we find robust training to beat dataset augmentation in every evaluation we performed; we find that a state-of-the-art architecture (code2seq) for models of code is harder to make robust than a simpler baseline; additionally, we find code2seq to have surprising weaknesses not present in our simpler baseline model; finally, we find that robust models perform better against unseen data from different sources (as one might hope)-however, we also find that robust models are not clearly better in the cross-language transfer task. To the best of our knowledge, we are the first to study the interplay between robustness of models of code and the domain-adaptation and cross-language- transfer tasks.