On the Reproducibility of Software Defect Datasets

On the Reproducibility of Software Defect Datasets
复制标题

DOI:
10.1109/icse48619.2023.00195
复制
发表时间:
2023-05
期刊:
2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE)
影响因子:
--
通讯作者:
Hao-Nan Zhu;Cindy Rubio-González
Hao-Nan Zhu;Cindy Rubio-González
中科院分区:
其他
文献类型:
--
作者:
Hao-Nan Zhu;Cindy Rubio-González

文献摘要

被引文献

相似文献

软件缺陷数据集对于促进故障定位、测试生成和自动程序修复等领域的技术评估和比较至关重要。然而,软件缺陷工件的可再现性并不能免于破损。在本文中,我们进行了软件缺陷工件的再现性的研究。首先,我们研究了五个最先进的Java缺陷数据集。尽管数据集维护者应用了多种策略来确保可重复性,但所有数据集都容易损坏。其次,我们进行了一个案例研究,在这个案例中,我们系统地测试了1,795个软件工件在13个月内的可重复性。我们发现,62.6%的文物至少打破一次,15.3%的文物打破多次。我们手动调查损坏的根本原因并手工制作10个补丁,这些补丁自动应用于2,948个修复中的1,055个不同工件。基于根本原因的性质,我们提出了自动依赖缓存和工件隔离,以防止进一步的破坏。特别是,我们表明,隔离工件以消除外部依赖性将再现性提高到95%或更高,这与最可靠的手动策展数据集所表现出的再现性水平相当。
Software defect datasets are crucial to facilitating the evaluation and comparison of techniques in fields such as fault localization, test generation, and automated program repair. However, the reproducibility of software defect artifacts is not immune to breakage. In this paper, we conduct a study on the reproducibility of software defect artifacts. First, we study five state-of-the-art Java defect datasets. Despite the multiple strategies applied by dataset maintainers to ensure reproducibility, all datasets are prone to breakages. Second, we conduct a case study in which we systematically test the reproducibility of 1,795 software artifacts during a 13-month period. We find that 62.6% of the artifacts break at least once, and 15.3% artifacts break multiple times. We manually investigate the root causes of breakages and handcraft 10 patches, which are automatically applied to 1,055 distinct artifacts in 2,948 fixes. Based on the nature of the root causes, we propose automated dependency caching and artifact isolation to prevent further breakage. In particular, we show that isolating artifacts to eliminate external dependencies increases reproducibility to 95% or higher, which is on par with the level of reproducibility exhibited by the most reliable manually curated dataset.