DebtFree: minimizing labeling cost in self-admitted technical debt identification using semi-supervised learning

DebtFree: minimizing labeling cost in self-admitted technical debt identification using semi-supervised learning
复制标题

DOI:
10.1007/s10664-022-10121-w
复制
发表时间:
2022-01
影响因子:
4.1
通讯作者:
Huy Tu;T. Menzies
Huy Tu;T. Menzies
中科院分区:
计算机科学2区
文献类型:
--
作者:
Huy Tu;T. Menzies

文献摘要

相似文献

跟踪和管理自我承认的技术债务(SATD)对于维护一个健康的软件项目非常重要。目前主动学习的SATD识别工具需要人工检查24%的测试评论,平均达到90%的召回率。在所有的测试评论中,大约5%是SATD。然后,人类专家需要阅读几乎五倍的SATD评论,这表明该工具的效率低下。此外,人类专家仍然容易出错:以前工作中95%的假阳性标签实际上是真阳性。为了解决上述问题,我们提出了DebtFree,一个基于无监督学习的双模式框架,用于识别SATD。在模式1中,当现有的训练数据未被标记时,DebtFree从一个无监督的学习者开始,自动伪标记训练数据中的编程评论。相比之下,在模式2中,标签与相应的训练数据一起可用,DebtFree从一个预处理器开始,该预处理器从测试数据集中识别出高度倾向的SATD。然后,我们的机器学习模型被用来帮助人类专家手动识别剩余的SATD。我们对10个软件项目的实验表明,这两个模型在统计上显着提高了最先进的自动化和半自动化模型的有效性。具体来说,DebtFree可以在模式1(未标记的训练数据)中减少99%的标记工作,在模式2(标记的训练数据)中减少63%,同时将当前主动学习者的F1相对提高到几乎100%。
Keeping track of and managing Self-Admitted Technical Debts (SATDs) is important for maintaining a healthy software project. Current active-learning SATD recognition tool involves manual inspection of 24% of the test comments on average to reach 90% of the recall. Among all the test comments, about 5% are SATDs. The human experts are then required to read almost a quintuple of the SATD comments which indicates the inefficiency of the tool. Plus, human experts are still prone to error: 95% of the false-positive labels from previous work were actually true positives. To solve the above problems, we propose DebtFree, a two-mode framework based on unsupervised learning for identifying SATDs. In mode1, when the existing training data is unlabeled, DebtFree starts with an unsupervised learner to automatically pseudo-label the programming comments in the training data. In contrast, in mode2 where labels are available with the corresponding training data, DebtFree starts with a pre-processor that identifies the highly prone SATDs from the test dataset. Then, our machine learning model is employed to assist human experts in manually identifying the remaining SATDs. Our experiments on 10 software projects show that both models yield statistically significant improvement in effectiveness over the state-of-the-art automated and semi-automated models. Specifically, DebtFree can reduce the labeling effort by 99% in mode1 (unlabeled training data), and up to 63%in mode2 (labeled training data) while improving the current active learner’s F1 relatively to almost 100%.