CCFinder: A multilinguistic token-based code clone detection system for large scale source code

CCFinder: A multilinguistic token-based code clone detection system for large scale source code
复制标题

DOI:
10.1109/tse.2002.1019480
复制
发表时间:
2002-07-01
影响因子:
7.4
通讯作者:
Inoue, K
Inoue, K
中科院分区:
计算机科学1区
文献类型:
--
作者:
Kamiya, T;Kusumoto, S;Inoue, K

文献摘要

被引文献

相似文献

代码克隆是源文件中与其他部分相同或相似的代码片段。由于人们认为代码克隆会降低软件的可维护性,因此已经提出了多种代码克隆检测技术和工具。本文提出了一种新的克隆检测技术,它包括对输入源文本的转换以及逐标记比较。通过使用几种有用的优化技术来实现它,我们开发了一个名为CCFinder的工具,它可以提取C、C++、Java、COBOL和其他源文件中的代码克隆。同时,还开发了用于代码克隆的度量指标。为了评估CCFinder和这些度量指标的有效性,我们进行了几个案例研究,将新工具应用于JDK、FreeBSD、NetBSD、Linux和许多其他系统的源代码。结果表明,CCFinder有效地找到了克隆,并且这些度量指标能够有效地识别系统的特征。此外,我们还将所提出的技术与其他克隆检测技术进行了比较。
A code clone is a code portion in source files that is identical or similar to another. Since code clones are believed to reduce the maintainability of software, several code clone detection techniques and tools have been proposed. This paper proposes a new clone detection technique, which consists of the transformation of input source text and a token-by-token comparison. For its implementation with several useful optimization techniques, we have developed a tool, named CCFinder, which extracts code clones in C, C++, Java, COBOL, and other source files. As well, metrics for the code clones have been developed. In order to evaluate the usefulness of CCFinder and metrics, we conducted several case studies where we applied the new tool to the source code of JDK, FreeBSD, NetBSD, Linux, and many other systems. As a result, CCFinder has effectively found clones and the metrics have been able to effectively identify the characteristics of the systems. In addition, we have compared the proposed technique with other clone detection techniques.