Neural Detection of Semantic Code Clones Via Tree-Based Convolution

Neural Detection of Semantic Code Clones Via Tree-Based Convolution
复制标题

DOI:
10.1109/icpc.2019.00021
复制
发表时间:
2019-05
期刊:
2019 IEEE/ACM 27th International Conference on Program Comprehension (ICPC)
影响因子:
--
通讯作者:
Hao Yu;Wing Lam;Long Chen;Ge Li;Tao Xie;Qianxiang Wang
Hao Yu;Wing Lam;Long Chen;Ge Li;Tao Xie;Qianxiang Wang
中科院分区:
其他
文献类型:
--
作者:
Hao Yu;Wing Lam;Long Chen;Ge Li;Tao Xie;Qianxiang Wang

文献摘要

被引文献

相似文献

代码克隆是类似的代码片段,它们共享相同的语义,但在语法上可能有不同程度的不同。检测代码克隆有助于降低软件维护成本并防止故障。在过去的二十年里,人们提出了各种检测代码克隆的方法,但很少有方法能够检测语义克隆,即具有不同语法的代码克隆。最近的研究尝试采用深度学习来检测代码克隆,例如在抽象语法树 (AST) 上使用基于树的 LSTM。然而,它没有充分利用代码片段的结构信息,从而限制了其克隆检测能力。为了充分发挥深度学习检测代码克隆的能力,我们提出了一种新方法,通过从 AST 捕获代码片段的结构信息和从代码标记中捕获词汇信息,使用基于树的卷积来检测语义克隆。此外,我们的方法解决了源代码具有无限的标记和模型词汇表的限制,因此在处理看不见的标记时,利用代码标记中的词汇信息通常是无效的。特别是,我们提出了一种称为位置感知字符嵌入(PACE)的新嵌入技术,该技术本质上将任何标记视为字符热嵌入的位置加权组合。我们的实验结果表明,我们的方法大大优于现有的最先进方法,在两个流行的代码克隆基准(OJClone 和 BigCloneBench)上的 F1 分数分别增加了 0.42 和 0.15,同时计算效率更高。我们的实验结果还表明,当代码克隆包含看不见的标记时,PACE 使我们的方法变得更加有效。
Code clones are similar code fragments that share the same semantics but may differ syntactically to various degrees. Detecting code clones helps reduce the cost of software maintenance and prevent faults. Various approaches of detecting code clones have been proposed over the last two decades, but few of them can detect semantic clones, i.e., code clones with dissimilar syntax. Recent research has attempted to adopt deep learning for detecting code clones, such as using tree-based LSTM over Abstract Syntax Tree (AST). However, it does not fully leverage the structural information of code fragments, thereby limiting its clone-detection capability. To fully unleash the power of deep learning for detecting code clones, we propose a new approach that uses tree-based convolution to detect semantic clones, by capturing both the structural information of a code fragment from its AST and lexical information from code tokens. Additionally, our approach addresses the limitation that source code has an unlimited vocabulary of tokens and models, and thus exploiting lexical information from code tokens is often ineffective when dealing with unseen tokens. Particularly, we propose a new embedding technique called position-aware character embedding (PACE), which essentially treats any token as a position-weighted combination of character one-hot embeddings. Our experimental results show that our approach substantially outperforms an existing state-of-the-art approach with an increase of 0.42 and 0.15 in F1-score on two popular code-clone benchmarks (OJClone and BigCloneBench), respectively, while being more computationally efficient. Our experimental results also show that PACE enables our approach to be substantially more effective when code clones contain unseen tokens.