Towards Automated Classification of Code Review Feedback to Support Analytics

Towards Automated Classification of Code Review Feedback to Support Analytics
复制标题

DOI:
10.1109/esem56168.2023.10304851
复制
发表时间:
2023-07
期刊:
2023 ACM/IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM)
影响因子:
--
通讯作者:
Asif Kamal Turzo;Fahim Faysal;Ovi Poddar;Jaydeb Sarker;Anindya Iqbal;Amiangshu Bosu
Asif Kamal Turzo;Fahim Faysal;Ovi Poddar;Jaydeb Sarker;Anindya Iqbal;Amiangshu Bosu
中科院分区:
其他
文献类型:
--
作者:
Asif Kamal Turzo;Fahim Faysal;Ovi Poddar;Jaydeb Sarker;Anindya Iqbal;Amiangshu Bosu

文献摘要

相似文献

背景:由于改进代码审查(CR)的有效性是许多软件开发组织的优先事项,项目已经部署了CR分析平台来识别潜在的改进领域。确定的问题数量是衡量CR有效性的关键指标,但如果将所有问题放在同一个分类箱中,可能会产生误导。因此,在CR期间确定的问题的细粒度分类可以提供可操作的见解,以提高CR的有效性。尽管弗雷格南等人最近的一项工作提出了自动模型来对cr引起的更改进行分类,但我们已经注意到两个潜在的改进领域——i)对不会引起更改的注释进行分类,ii)将深度神经网络(DNN)与代码上下文结合使用以提高性能。目的:本研究旨在开发一种自动CR评论分类器,该分类器利用DNN模型实现比弗雷格南等人更可靠的性能。方法:使用1828条CR评论的人工标记数据集,我们训练并评估了基于监督学习的DNN模型,该模型利用代码上下文、评论文本和一组代码度量来将CR评论分类为Turzo和Bosu提出的五个高级类别之一。结果:基于我们对多种标记化方法组合的10倍交叉验证评估,我们发现使用CodeBERT的模型达到了59.3%的最佳准确率。我们的方法优于弗雷格南等人的方法,准确率提高了18.7%。结论:除了促进改进的CR分析之外,我们提出的模型对于开发人员确定代码评审反馈的优先级和选择评审人员也很有用。
Background: As improving code review (CR) effectiveness is a priority for many software development organizations, projects have deployed CR analytics platforms to identify potential improvement areas. The number of issues identified, which is a crucial metric to measure CR effectiveness, can be misleading if all issues are placed in the same bin. Therefore, a finer-grained classification of issues identified during CRs can provide actionable insights to improve CR effectiveness. Although a recent work by Fregnan et al. proposed automated models to classify CR-induced changes, we have noticed two potential improvement areas – i) classifying comments that do not induce changes and ii) using deep neural networks (DNN) in conjunction with code context to improve performances. Aims: This study aims to develop an automated CR comment classifier that leverages DNN models to achieve a more reliable performance than Fregnan et al. Method: Using a manually labeled dataset of 1,828 CR comments, we trained and evaluated supervised learning-based DNN models leveraging code context, comment text, and a set of code metrics to classify CR comments into one of the five high-level categories proposed by Turzo and Bosu. Results: Based on our 10-fold cross-validation-based evaluations of multiple combinations of tokenization approaches, we found a model using CodeBERT achieving the best accuracy of 59.3%. Our approach outperforms Fregnan et al.'s approach by achieving 18.7% higher accuracy. Conclusion: In addition to facilitating improved CR analytics, our proposed model can be useful for developers in prioritizing code review feedback and selecting reviewers.