Local-based active classification of test report to assist crowdsourced testing

Local-based active classification of test report to assist crowdsourced testing
复制标题

DOI:
10.1145/2970276.2970300
复制
发表时间:
2016-08
期刊:
2016 31st IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子:
--
通讯作者:
Junjie Wang;Song Wang;Qiang Cui;Qing Wang
Junjie Wang;Song Wang;Qiang Cui;Qing Wang
中科院分区:
其他
文献类型:
--
作者:
Junjie Wang;Song Wang;Qiang Cui;Qing Wang

文献摘要

被引文献

相似文献

在众包测试中,一项重要的任务是从众包工作人员提交的大量测试报告中识别出真正揭示故障的测试报告。解决这个问题的大多数现有方法都使用了有监督的机器学习技术,这通常需要用户手动标记大量的训练数据。这样的过程既耗时又费力。因此,减少人工贴标签的繁重负担,同时仍然能够获得良好的性能是至关重要的。主动学习是解决这一挑战的一种潜在技术,其目的是用尽可能少的标签数据训练一个好的分类器。然而,我们对真实工业数据的观察表明,现有的主动学习方法在众包测试数据上会产生糟糕和不稳定的性能。我们分析了深层次的原因,发现数据集具有显著的局部偏差。针对上述问题,我们提出了基于局部的主动分类方法(LOAF)来从众包测试报告中对真实故障进行分类。Loaf推荐一小部分局部邻域内信息最丰富的实例,并询问用户它们的标签,然后学习基于局部邻域的分类器。我们对来自中国最大的众包测试平台之一的34个商业项目的14,609份测试报告进行了评估,结果表明,我们提出的LOAF可以产生良好的结果。此外,它的性能甚至优于现有的基于大量已标记历史数据的监督学习方法。此外,我们还实施了我们的方法,并使用真实世界的案例研究来评估其有用性。测试人员的反馈证明了该方法的实用价值。
In crowdsourced testing, an important task is to identify the test reports that actually reveal fault - true fault, from the large number of test reports submitted by crowd workers. Most existing approaches towards this problem utilized supervised machine learning techniques, which often require users to manually label a large amount of training data. Such process is time-consuming and labor-intensive. Thus, reducing the onerous burden of manual labeling while still being able to achieve good performance is crucial. Active learning is one potential technique to address this challenge, which aims at training a good classifier with as few labeled data as possible. Nevertheless, our observation on real industrial data reveals that existing active learning approaches generate poor and unstable performances on crowdsourced testing data. We analyze the deep reason and find that the dataset has significant local biases. To address the above problems, we propose LOcal-based Active ClassiFication (LOAF) to classify true fault from crowdsourced test reports. LOAF recommends a small portion of instances which are most informative within local neighborhood, and asks user their labels, then learns classifiers based on local neighborhood. Our evaluation on 14,609 test reports of 34 commercial projects from one of the Chinese largest crowdsourced testing platforms shows that our proposed LOAF can generate promising results. In addition, its performance is even better than existing supervised learning approaches which built on large amounts of labelled historical data. Moreover, we also implement our approach and evaluate its usefulness using real-world case studies. The feedbacks from testers demonstrate its practical value.