Towards Privacy Preserving Cross Project Defect Prediction with Federated Learning

Towards Privacy Preserving Cross Project Defect Prediction with Federated Learning
复制标题

DOI:
10.1109/saner56733.2023.00052
复制
发表时间:
2023-03
期刊:
2023 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER)
影响因子:
--
通讯作者:
Hiroki Yamamoto;Dong Wang;Gopi Krishnan Rajbahadur;Masanari Kondo;Yasutaka Kamei;Naoyasu Ubayashi
Hiroki Yamamoto;Dong Wang;Gopi Krishnan Rajbahadur;Masanari Kondo;Yasutaka Kamei;Naoyasu Ubayashi
中科院分区:
其他
文献类型:
--
作者:
Hiroki Yamamoto;Dong Wang;Gopi Krishnan Rajbahadur;Masanari Kondo;Yasutaka Kamei;Naoyasu Ubayashi

文献摘要

相似文献

缺陷预测模型可以预测软件项目中的缺陷,许多研究人员研究缺陷预测模型来辅助软件开发中的调试工作。近年来,跨项目缺陷预测(CPDP)受到了越来越多的关注,它是在没有足够的数据来构建缺陷预测模型的情况下,使用从其他项目的数据中学习的缺陷预测模型来预测项目中的缺陷。由于CPDP使用其他项目的数据,数据隐私保护是最重要的问题之一。然而,以前的CPDP研究仍然需要项目之间的数据共享来培训模型,并且没有充分考虑保护项目的机密性。为了解决这一问题,我们提出了一种采用联合学习的CPDP模型FLR,这是一种不需要数据共享的分布式机器学习方法。我们评估FLR,使用25个项目,以调查其有效性和特征解释。我们的关键结果表明,首先,FLR比现有的隐私保护方法(即LACE2)具有更好的性能。同时,该方法的性能与传统方法(如监督学习和非监督学习)相当。其次,解释分析的结果表明,尺度相关特征对FLR的预测性能有共同的影响。此外,进一步的研究表明,联合学习的参数(例如,学习率和客户端数量)也对性能起到作用。本研究作为第一步,证实了联合学习在CPDP中用于隐私保护的可行性,并为未来将其他机器学习模型应用于联合学习的研究奠定了基础。
Defect prediction models can predict defects in software projects, and many researchers study defect prediction models to assist debugging efforts in software development. In recent years, there has been growing interest in Cross Project Defect Prediction (CPDP), which predicts defects in a project using a defect prediction model learned from other projects’ data when there is insufficient data to construct a defect prediction model. Since CPDP uses other projects’ data, data privacy preservation is one of the most significant issues. However, prior CPDP studies still require data sharing among projects to train models, and do not fully consider protecting project confidentiality. To address this, we propose a CPDP model FLR employing federated learning, a distributed machine learning approach that does not require data sharing. We evaluate FLR, using 25 projects, to investigate its effectiveness and feature interpretation. Our key results show that first, FLR outperforms the existing privacy-preserving methods (i.e., LACE2). Meanwhile, the performance is relatively comparable to the conventional methods (e.g., supervised and unsupervised learning). Second, the results of the interpretation analysis show that scale-related features have a common effect on the prediction performance of the FLR. In addition, further insights demonstrate that parameters of federated learning (e.g., learning rates and the number of clients) also play a role in the performance. This study is served as a first step to confirm the feasibility of the employment of federated learning in CPDP to ensure privacy preservation and lays the groundwork for future research on applying other machine learning models to federated learning.