REPENT: Analyzing the Nature of Identifier Renamings

REPENT: Analyzing the Nature of Identifier Renamings
复制标题

DOI:
10.1109/tse.2014.2312942
复制
发表时间:
2014-05
影响因子:
7.4
通讯作者:
Venera Arnaoudova;L. Eshkevari;M. D. Penta;Rocco Oliveto;G. Antoniol;Yann-Gaël Guéhéneuc
Venera Arnaoudova;L. Eshkevari;M. D. Penta;Rocco Oliveto;G. Antoniol;Yann-Gaël Guéhéneuc
中科院分区:
计算机科学1区
文献类型:
--
作者:
Venera Arnaoudova;L. Eshkevari;M. D. Penta;Rocco Oliveto;G. Antoniol;Yann-Gaël Guéhéneuc

文献摘要

被引文献

相似文献

源代码词汇在软件质量中起着至关重要的作用:糟糕的词汇会导致较差的可理解性,甚至增加软件的错误倾向。由于这个原因,重命名一个程序实体,即改变实体标识符,是软件发展过程中的一个重要活动。当开发人员觉得一个实体的名称(不再)与其功能一致时,或者当这样的名称可能具有误导性时,他们就会重命名。我们对71名开发人员进行的一项调查表明,39%的人执行重命名,从每周几次到几乎每天一次,92%的参与者认为重命名并不简单。然而,尽管与重命名相关的成本很高,但重命名很少被记录下来——例如,在我们研究的五个程序中,只有不到1%的重命名被记录下来。这解释了为什么参与者基本上同意自动记录重命名的有用性。在本文中,我们提出了重命名程序实体(REanaming Program ENTities,悔改),这是一种在源代码中自动检测文档和分类标识符重命名的方法。忏悔基于源代码差异和数据流分析的组合来检测重命名。使用一组自然语言工具,懊悔将重命名分类到我们定义的分类法的不同维度中。使用文档化的重命名,开发人员将能够,例如,查找作为公共API的一部分的方法(因为它们影响客户端应用程序),或者查找经历高风险重命名的实体的名称和实现之间的不一致(例如,朝着相反的含义)。我们在五个开源Java程序的进化史上评估了忏悔的准确性和完整性。研究表明,准确率为88%,召回率为92%。此外,我们报告了一项探索性研究,调查和讨论了如何根据我们的分类法在五个程序中重新命名标识符。
Source code lexicon plays a paramount role in software quality: poor lexicon can lead to poor comprehensibility and even increase software fault-proneness. For this reason, renaming a program entity, i.e., altering the entity identifier, is an important activity during software evolution. Developers rename when they feel that the name of an entity is not (anymore) consistent with its functionality, or when such a name may be misleading. A survey that we performed with 71 developers suggests that 39 percent perform renaming from a few times per week to almost every day and that 92 percent of the participants consider that renaming is not straightforward. However, despite the cost that is associated with renaming, renamings are seldom if ever documented-for example, less than 1 percent of the renamings in the five programs that we studied. This explains why participants largely agree on the usefulness of automatically documenting renamings. In this paper we propose REanaming Program ENTities (REPENT), an approach to automatically document-detect and classify-identifier renamings in source code. REPENT detects renamings based on a combination of source code differencing and data flow analyses. Using a set of natural language tools, REPENT classifies renamings into the different dimensions of a taxonomy that we defined. Using the documented renamings, developers will be able to, for example, look up methods that are part of the public API (as they impact client applications), or look for inconsistencies between the name and the implementation of an entity that underwent a high risk renaming (e.g., towards the opposite meaning). We evaluate the accuracy and completeness of REPENT on the evolution history of five open-source Java programs. The study indicates a precision of 88 percent and a recall of 92 percent. In addition, we report an exploratory study investigating and discussing how identifiers are renamed in the five programs, according to our taxonomy.