Similarity of Source Code in the Presence of Pervasive Modifications

Similarity of Source Code in the Presence of Pervasive Modifications
复制标题

存在普遍修改的源代码的相似性

DOI:
--
复制
发表时间:
2016
期刊:
IEEE Working Conference on Source Code Analysis and Manipulation
影响因子:
--
通讯作者:
D. Clark
D. Clark
中科院分区:
--
文献类型:
--
作者:
Chaiyong Ragkhitwetsagul;J. Krinke;D. Clark

文献摘要

被引文献

相似文献

用于检测代码克隆、代码抄袭和代码重用的源代码分析面临着普遍的代码修改问题,即可能产生全局影响的转换。我们将 30 种相似性检测技术和工具与普遍的代码修改进行比较。我们使用 Java 源代码的两个实验场景来评估这些工具。这些是(1)使用源代码和字节码混淆工具创建的普遍修改,以及(2)通过使用不同的反编译器进行编译和反编译的源代码规范化。我们的实验结果表明,高度专业化的源代码相似性检测技术和工具可以比更通用的文本相似性度量表现得更好。我们的研究有力地验证了编译/反编译作为规范化技术的使用。它的使用将六种工具的错误分类减少到零。这项广泛、彻底的研究是现有规模最大的研究,对于未来源代码中相似性检测的用户来说可能是一个宝贵的指南。
Source code analysis to detect code cloning, code plagiarism, and code reuse suffers from the problem of pervasive code modifications, i.e. transformations that may have a global effect. We compare 30 similarity detection techniques and tools against pervasive code modifications. We evaluate the tools using two experimental scenarios for Java source code. These are (1) pervasive modifications created with tools for source code and bytecode obfuscation and (2) source code normalisation through compilation and decompilation using different decompilers. Our experimental results show that highly specialised source code similarity detection techniques and tools can perform better than more general, textual similarity measures. Our study strongly validates the use of compilation/decompilation as a normalisation technique. Its use reduced false classifications to zero for six of the tools. This broad, thorough study is the largest in existence and potentially an invaluable guide for future users of similarity detection in source code.