Just‐in‐time identification for cross‐project correlated issues

Just‐in‐time identification for cross‐project correlated issues
复制标题

DOI:
10.1002/smr.2637
复制
发表时间:
2023-12
期刊:
Journal of Software: Evolution and Process
影响因子:
--
通讯作者:
Hao Ren;Yanhui Li;Lin Chen;Yulu Cao;Xiaowei Zhang;Changhai Nie
Hao Ren;Yanhui Li;Lin Chen;Yulu Cao;Xiaowei Zhang;Changhai Nie
中科院分区:
其他
文献类型:
--
作者:
Hao Ren;Yanhui Li;Lin Chen;Yulu Cao;Xiaowei Zhang;Changhai Nie

文献摘要

相似文献

问题跟踪系统在软件开发中非常流行,它可以帮助开发人员提交和讨论问题,以解决软件项目中的开发问题。以前的大多数研究都是为了分析项目中的问题关系,例如推荐类似或重复的错误问题。然而,随着多项目联合开发的普及,沿着,许多问题是跨项目相关的(CPC),即一个问题与不同项目中的另一个问题相关联。当开发人员遇到CPC问题时,它可能主要增加解决这些问题的难度,因为他们不仅需要来自其项目的信息,而且还需要开发人员不熟悉的其他相关项目的信息。尽早识别CPC问题是管理者和开发者分配软件维护资源和估计解决问题所需努力的基本挑战。本文提出了两组11个问题度量,用于描述文本摘要和报告者的活动,这些度量可以在问题报告后立即提取。我们使用这11个问题指标来构建即时(JIT)预测模型,以识别CPC问题。为了评估CPC问题预测模型的效果,我们对16个开源数据科学和深度学习项目进行了实验,并将我们的预测模型与基于文本特征的两个基线模型进行了比较(即,词频-逆文档频率[TF-IDF]和词嵌入),这是以前关于问题预测的研究中常用的方法。结果表明,基于问题度量的JIT预测模型在两个评价指标下,即MCC和F1下,显著提高了CPC问题预测的性能。此外,我们发现该预测模型更适合于开源生态系统中的大型复杂核心项目。
Issue tracking systems are now prevalent in software development, which would help developers submit and discuss issues to solve development problems on software projects. Most previous studies have been conducted to analyze issue relations within projects, such as recommending similar or duplicate bug issues. However, along with the popularization of co‐developing through multiple projects, many issues are cross‐project correlated (CPC), that is, one issue is associated with another issue in a different project. When developers meet with CPC issues, it may primarily increase the difficulties of solving them because they need information from not only their projects but also other related projects that developers are not familiar with. Identifying a CPC issue as early as possible is a fundamental challenge for both managers and developers to allocate the resources for software maintenance and estimate the effort to solve it. This paper proposes 11 issue metrics of two groups to describe textual summary and reporters' activity, which can be extracted just after the issue was reported. We employ these 11 issue metrics to construct just‐in‐time (JIT) prediction models to identify CPC issues. To evaluate the effect of CPC issue prediction models, we conduct experiments on 16 open‐source data science and deep learning projects and compare our prediction model with two baseline models based on textual features (i.e., Term Frequency‐Inverse Document Frequency [TF‐IDF] and Word Embedding), which are commonly adopted by previous studies on issue prediction. The results show that the JIT prediction model based on issue metrics has significantly improved the performance of CPC issue prediction under two evaluation indicators, Matthew's correlation coefficient (MCC) and F1. In addition, we find that the prediction model is more suitable for large‐scale complex core projects in the open‐source ecosystem.