An investigation of online and offline learning models for online Just-in-Time Software Defect Prediction

An investigation of online and offline learning models for online Just-in-Time Software Defect Prediction
复制标题

DOI:
10.1007/s10664-023-10335-6
复制
发表时间:
2023-09
影响因子:
4.1
通讯作者:
George G. Cabral;Leandro L. Minku;Adriano L. I. Oliveira;Dinaldo A. Pessoa;Sadia Tabassum
George G. Cabral;Leandro L. Minku;Adriano L. I. Oliveira;Dinaldo A. Pessoa;Sadia Tabassum
中科院分区:
计算机科学2区
文献类型:
--
作者:
George G. Cabral;Leandro L. Minku;Adriano L. I. Oliveira;Dinaldo A. Pessoa;Sadia Tabassum

文献摘要

相似文献

即时软件缺陷预测 (JIT-SDP) 在在线场景中运行,随着时间的推移会收到额外的训练数据。现有的在线 JIT-SDP 研究使用在线 Oza 集成学习方法,以 Hoeffding 树作为基础学习器,在这种情况下随着时间的推移学习和更新 JIT-SDP 模型。然而,目前尚不清楚这些方法与适用于在线场景的离线学习方法相比如何,以及使用任何其他在线或离线基础学习器将如何影响在线 JIT-SDP 的预测性能和计算成本。因此,我们提出了一种称为批量过采样率提升(BORB)的新方法,它能够在在线 JIT-SDP 场景中使用离线基础学习器。基于 10 个开源项目,我们在项目内和跨项目在线 JIT-SDP 场景中,对具有 5 个不同基础学习器的 BORB 以及具有 4 个不同基础学习器的现有在线方法过采样率提升进行了综合评估。结果表明,与我们研究中考虑的性能最佳的在线学习方法相比,离线学习可以带来更好的预测性能,但计算成本更高。跨项目数据有助于提高离线和在线学习的预测性能,尤其是在线学习。
Just-in-Time Software Defect Prediction (JIT-SDP) operates in an online scenario where additional training data is received over time. Existing online JIT-SDP studies used online Oza ensemble learning methods with Hoeffding Trees as base learners to learn and update JIT-SDP models over time in this scenario. However, it is unknown how these approaches compare against offline learning approaches adapted to operate in online scenarios, and how the use of any other online or offline base learners would affect online JIT-SDP in terms of predictive performance and computational cost. We therefore propose a new approach called Batch Oversampling Rate Boosting (BORB) that is able to use offline base learners in an online JIT-SDP scenario. Based on 10 open source projects, we provide a comprehensive evaluation of BORB with 5 different base learners and the existing online approach Oversampling Rate Boosting with 4 different base learners, both in within-project and cross-project online JIT-SDP scenarios. The results show that offline learning can lead to better predictive performance than the top performing online learning approaches considered in our study, at a higher computational cost. Cross-project data was helpful to improve predictive performance both for offline and online learning, but especially for online learning.