Code Similarities Beyond Copy & Paste

Code Similarities Beyond Copy & Paste
复制标题

超越复制的代码相似之处

DOI:
10.1109/csmr.2010.33
复制
发表时间:
2010
期刊:
2010 14th European Conference on Software Maintenance and Reengineering
影响因子:
--
通讯作者:
B. Hummel
B. Hummel
中科院分区:
--
文献类型:
--
作者:
Elmar Jürgens;F. Deißenböck;B. Hummel

文献摘要

被引文献

相似文献

冗余的源代码阻碍了软件维护,因为更新必须在多个地方执行。这与冗余是由复制和粘贴还是由行为相似的代码的独立开发所创建无关。现有的克隆检测工具可以成功地发现语法相似的冗余代码。因此,它们可以很好地处理复制和粘贴所产生的冗余。但是:独立起源的行为相似代码在句法上有多相似?本文介绍了一个对照实验的结果表明,独立起源的行为相似的代码是极不可能是句法相似的。事实上,它在语法上是如此不同,以至于现有的克隆检测方法不能识别超过1%的这种冗余。这是不幸的,因为对开源软件的人工检查表明,独立起源的行为相似的代码在实践中确实存在,并且确实存在维护问题。
Redundant source code hinders software maintenance, since updates have to be performed in multiple places. This holds independent of whether redundancy was created by copy&paste or by independent development of behaviorally similar code. Existing clone detection tools successfully discover syntactically similar redundant code. They thus work well for redundancy that has been created by copy&paste. But: how syntactically similar is behaviorally similar code of independent origin? This paper presents the results of a controlled experiment that demonstrates that behaviorally similar code of independent origin is highly unlikely to be syntactically similar. In fact, it is so syntactically different, that existing clone detection approaches cannot identify more than 1% of such redundancy. This is unfortunate, as manual inspections of open source software indicate that behaviorally similar code of independent origin does exist in practice and does present problems to maintenance.