FRUGAL: Unlocking Semi-Supervised Learning for Software Analytics

FRUGAL: Unlocking Semi-Supervised Learning for Software Analytics
复制标题

DOI:
10.1109/ase51524.2021.9678617
复制
发表时间:
2021-08
期刊:
2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE)
影响因子:
--
通讯作者:
Huy Tu;T. Menzies
Huy Tu;T. Menzies
中科院分区:
其他
文献类型:
--
作者:
Huy Tu;T. Menzies

文献摘要

被引文献

相似文献

标准的软件分析通常涉及大量带有标签的数据,以便以可接受的性能委托模型。然而,先前的工作表明,这样的要求可能是昂贵的,需要几周的时间来标记数千个提交,并且在遍历新的研究问题和领域时并不总是可用。无监督学习是学习未标记数据中隐藏模式的一个很有前途的方向,它只在缺陷预测中得到了广泛的研究。然而,无监督学习本身可能是无效的,并且尚未在其他领域进行探索(例如,静态分析和问题关闭时间)。受该文献空白和技术限制的激励,我们提出了FRUGAL,一种基于简单优化方案的调谐半监督方法,其不需要复杂的(例如,深度学习器)和昂贵的(例如,100%手动标记数据)方法。FRUGAL优化了无监督学习器的配置(通过一个简单的网格搜索),同时验证了我们的设计决策的标签只有2.5%的数据之前prediction.As本文的实验所示,FRUGAL优于国家的最先进的可采用的静态代码警告识别器和问题关闭的时间预测,同时减少了40倍的标签成本(从100%到2.5%)。因此,我们断言,FRUGAL可以节省相当大的努力,在数据标签,特别是在验证以前的工作或研究新的problems.Based上这项工作,我们建议,复杂和昂贵的方法的支持者应该始终基线这样的方法对更简单,更便宜的替代品。例如,像FRUGAL这样的半监督学习器可以作为最先进的软件分析的基线。
Standard software analytics often involves having a large amount of data with labels in order to commission models with acceptable performance. However, prior work has shown that such requirements can be expensive, taking several weeks to label thousands of commits, and not always available when traversing new research problems and domains. Unsupervised Learning is a promising direction to learn hidden patterns within unlabelled data, which has only been extensively studied in defect prediction. Nevertheless, unsupervised learning can be ineffective by itself and has not been explored in other domains (e.g., static analysis and issue close time).Motivated by this literature gap and technical limitations, we present FRUGAL, a tuned semi-supervised method that builds on a simple optimization scheme that does not require sophisticated (e.g., deep learners) and expensive (e.g., 100% manually labelled data) methods. FRUGAL optimizes the unsupervised learner’s configurations (via a simple grid search) while validating our design decision of labelling just 2.5% of the data before prediction.As shown by the experiments of this paper FRUGAL outperforms the state-of-the-art adoptable static code warning recognizer and issue closed time predictor, while reducing the cost of labelling by a factor of 40 (from 100% to 2.5%). Hence we assert that FRUGAL can save considerable effort in data labelling especially in validating prior work or researching new problems.Based on this work, we suggest that proponents of complex and expensive methods should always baseline such methods against simpler and cheaper alternatives. For instance, a semi-supervised learner like FRUGAL can serve as a baseline to the state-of-theart software analytics.