On the Impact of Programming Languages on Code Quality

On the Impact of Programming Languages on Code Quality
复制标题

论编程语言对代码质量的影响

DOI:
--
复制
发表时间:
2019
影响因子:
1.3
通讯作者:
J. Vitek
J. Vitek
中科院分区:
计算机科学2区
文献类型:
--
作者:
E. Berger;Celeste Hollenbeck;Petr Maj;O. Vitek;J. Vitek

文献摘要

被引文献

相似文献

在2014年的文章中,Ray,Posnett,Devanbu和Filkov声称在Github上托管的729个项目中发现了11种编程语言和软件缺陷之间具有统计学意义的关联。具体而言,他们的工作回答了与软件缺陷和编程语言有关的四个研究问题。借助作者提供的数据和代码,本文首先尝试对原始研究进行实验重复。由于缺少语言分类的代码和问题,重复仅部分成功。这项工作的第二部分重点是他们的主要主张,即错误和语言之间的关联,并对数据和Ray等人进行的统计建模步骤进行了完整,独立的重新分析。在2014年。这项重新分析发现了许多严重的缺陷,这些缺陷将语言数量减少,而与缺陷的关联减少到只有11个。此外,实际效应大小非常小。因此,这些结果破坏了原始研究的结论。纠正记录很重要,因为许多随后的作品都引用了2014年的文章,并在没有证据的情况下断言了针对给定任务的编程语言和软件缺陷的数量之间的因果关系。手头数据不支持因果关系;而且,我们认为,即使修复了我们发现的方法论缺陷,太多的偏见来源太多,希望对跨语言的错误率进行有意义的比较。
In a 2014 article, Ray, Posnett, Devanbu, and Filkov claimed to have uncovered a statistically significant association between 11 programming languages and software defects in 729 projects hosted on GitHub. Specifically, their work answered four research questions relating to software defects and programming languages. With data and code provided by the authors, the present article first attempts to conduct an experimental repetition of the original study. The repetition is only partially successful, due to missing code and issues with the classification of languages. The second part of this work focuses on their main claim, the association between bugs and languages, and performs a complete, independent reanalysis of the data and of the statistical modeling steps undertaken by Ray et al. in 2014. This reanalysis uncovers a number of serious flaws that reduce the number of languages with an association with defects down from 11 to only 4. Moreover, the practical effect size is exceedingly small. These results thus undermine the conclusions of the original study. Correcting the record is important, as many subsequent works have cited the 2014 article and have asserted, without evidence, a causal link between the choice of programming language for a given task and the number of software defects. Causation is not supported by the data at hand; and, in our opinion, even after fixing the methodological flaws we uncovered, too many unaccounted sources of bias remain to hope for a meaningful comparison of bug rates across languages.