Learning a Metric for Code Readability

Learning a Metric for Code Readability
复制标题

DOI:
10.1109/tse.2009.70
复制
发表时间:
2010-07-01
影响因子:
7.4
通讯作者:
Weimer, Westley R.
Weimer, Westley R.
中科院分区:
计算机科学1区
文献类型:
--
作者:
Buse, Raymond P. L.;Weimer, Westley R.

文献摘要

被引文献

相似文献

在本文中,我们探讨了代码可读性的概念,并调查其与软件质量的关系。通过从120个人类注释者收集的数据,我们得出了一组简单的本地代码功能和人类可读性概念之间的关联。使用这些功能,我们构建了一个自动化的可读性度量,并表明它可以有80%的有效性,平均来说,在预测可读性判断方面比人类更好。此外,我们表明,这个度量与软件质量的三个措施:代码更改,自动缺陷报告和缺陷日志消息密切相关。我们在超过220万行代码上测量这些相关性,并在选定项目的许多版本上纵向测量。最后,我们讨论了这项研究对编程语言设计和工程实践的影响。例如,我们的数据表明,评论本身并不比简单的空白行对可读性的本地判断更重要。
In this paper, we explore the concept of code readability and investigate its relation to software quality. With data collected from 120 human annotators, we derive associations between a simple set of local code features and human notions of readability. Using those features, we construct an automated readability measure and show that it can be 80 percent effective and better than a human, on average, at predicting readability judgments. Furthermore, we show that this metric correlates strongly with three measures of software quality: code changes, automated defect reports, and defect log messages. We measure these correlations on over 2.2 million lines of code, as well as longitudinally, over many releases of selected projects. Finally, we discuss the implications of this study on programming language design and engineering practice. For example, our data suggest that comments, in and of themselves, are less important than simple blank lines to local judgments of readability.