Data Quality: Some Comments on the NASA Software Defect Datasets

Data Quality: Some Comments on the NASA Software Defect Datasets
复制标题

数据质量:对 NASA 软件缺陷数据集的一些评论

DOI:
10.1109/tse.2013.11
复制
发表时间:
2013-09-01
影响因子:
7.4
通讯作者:
Mair, Carolyn
Mair, Carolyn
中科院分区:
计算机科学1区
文献类型:
--
作者:
Shepperd, Martin;Song, Qinbao;Mair, Carolyn

文献摘要

被引文献

相似文献

背景——不言而喻的实证分析依赖于数据的质量。同样,复制依赖于准确的报告和使用相同而不是相似版本的数据集。近年来,人们对使用机器学习器将软件模块分为容易出现缺陷和不易出现缺陷的类别产生了很大的兴趣。公开可用的 NASA 数据集已被广泛用作这项研究的一部分。目的——这篇简短的文章调查了基于 NASA 缺陷数据集的已发表分析的有意义和可比性的程度。方法 - 我们分析了自 2007 年以来发表在《IEEE 软件工程汇刊》上的五项研究,这些研究利用了这些数据集,并比较了当前使用的数据集的两个版本。结果 - 我们发现两个版本的数据集之间存在重要差异,一个数据集中的值令人难以置信,并且数据集预处理的详细记录通常不足。结论 - 建议研究人员 1)指出他们使用的数据集的出处,2)足够详细地报告任何预处理以实现有意义的复制,3)在应用机器学习器之前投入精力理解数据。
Background-Self-evidently empirical analyses rely upon the quality of their data. Likewise, replications rely upon accurate reporting and using the same rather than similar versions of datasets. In recent years, there has been much interest in using machine learners to classify software modules into defect-prone and not defect-prone categories. The publicly available NASA datasets have been extensively used as part of this research. Objective-This short note investigates the extent to which published analyses based on the NASA defect datasets are meaningful and comparable. Method-We analyze the five studies published in the IEEE Transactions on Software Engineering since 2007 that have utilized these datasets and compare the two versions of the datasets currently in use. Results-We find important differences between the two versions of the datasets, implausible values in one dataset and generally insufficient detail documented on dataset preprocessing. Conclusions-It is recommended that researchers 1) indicate the provenance of the datasets they use, 2) report any preprocessing in sufficient detail to enable meaningful replication, and 3) invest effort in understanding the data prior to applying machine learners.