SLACC: Simion-based Language Agnostic Code Clones

SLACC: Simion-based Language Agnostic Code Clones
复制标题

DOI:
10.1145/3377811.3380407
复制
发表时间:
2020-02
期刊:
2020 IEEE/ACM 42nd International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
George Mathew;Chris Parnin;Kathryn T. Stolee
George Mathew;Chris Parnin;Kathryn T. Stolee
中科院分区:
其他
文献类型:
--
作者:
George Mathew;Chris Parnin;Kathryn T. Stolee

文献摘要

相似文献

成功的跨语言克隆检测可以使研究人员和开发人员能够创建强大的语言迁移工具,在掌握一种编程语言后促进学习其他编程语言,并在更广泛的代码库中促进代码片段的重用。然而,识别跨语言克隆对克隆检测问题提出了特殊的挑战。在任意语言之间缺乏公共底层表示意味着检测克隆需要以下解决方案之一:1)跨每种目标语言复制的静态分析框架,其注释跨所有语言匹配语言特征,或者2)基于运行时行为检测克隆的动态分析框架。在这项工作中,我们论证了后一种解决方案的可行性,即一种称为SLACC的跨语言克隆检测的动态分析方法。像以前的克隆检测技术一样,我们使用输入/输出行为来匹配克隆,尽管我们通过放大输入数量和覆盖更多数据类型克服了以前工作的限制;因此,实现了比以前尝试的更好的集群。由于集群是基于输入/输出行为生成的,因此SLACC支持跨语言克隆检测。作为额外的挑战,我们将目标对准静态类型化语言Java和动态类型化语言Python。与最新的Java克隆检测工具HitoshiIO相比,SLACC检索的集群数量是HitoshiIO的6倍,精度更高(86.7%比30.7%)。这是第一个对动态类型语言执行克隆检测的工作(精度=87.3%),也是第一个跨缺乏公共底层表示形式的语言执行克隆检测的工作(精度=94.1%)。它提供了迈向可伸缩语言迁移工具这一更大目标的第一步。
Successful cross-language clone detection could enable researchers and developers to create robust language migration tools, facilitate learning additional programming languages once one is mastered, and promote reuse of code snippets over a broader codebase. However, identifying cross-language clones presents special challenges to the clone detection problem. A lack of common underlying representation between arbitrary languages means detecting clones requires one of the following solutions: 1) a static analysis framework replicated across each targeted language with annotations matching language features across all languages, or 2) a dynamic analysis framework that detects clones based on runtime behavior. In this work, we demonstrate the feasibility of the latter solution, a dynamic analysis approach called SLACC for cross-language clone detection. Like prior clone detection techniques, we use input/output behavior to match clones, though we overcome limitations of prior work by amplifying the number of inputs and covering more data types; and as a result, achieve better clusters than prior attempts. Since clusters are generated based on input/output behavior, SLACC supports cross-language clone detection. As an added challenge, we target a static typed language, Java, and a dynamic typed language, Python. Compared to HitoshiIO, a recent clone detection tool for Java, SLACC retrieves 6 times as many clusters and has higher precision (86.7% vs. 30.7%). This is the first work to perform clone detection for dynamic typed languages (precision = 87.3%) and the first to perform clone detection across languages that lack a common underlying representation (precision = 94.1%). It provides a first step towards the larger goal of scalable language migration tools.