A novel online supervised hyperparameter tuning procedure applied to cross-company software effort estimation

A novel online supervised hyperparameter tuning procedure applied to cross-company software effort estimation
复制标题

DOI:
10.1007/s10664-019-09686-w
复制
发表时间:
2019-02
影响因子:
4.1
通讯作者:
Leandro L. Minku
Leandro L. Minku
中科院分区:
计算机科学2区
文献类型:
--
作者:
Leandro L. Minku

文献摘要

被引文献

相似文献

软件工作量估计是一个在线监督学习问题,随着时间的推移,新的培训项目可能会变得可用。在这种情况下,Dycom的跨公司(CC)方法可以大大减少培训所需的公司内(WC)项目数量,节省收集成本。但是,Dycom要求将CC项目拆分为子集。这些子集的数量和组成都会影响Dycom的预测性能。即使聚类方法可以用于自动创建CC子集,也没有用于在在线监督场景中随时间自动调整聚类数量的过程。本文提出了第一个程序。Dycom使用六个聚类方法和三个自动调整程序进行调查,以检查是否聚类自动调整可以创建良好的CC分裂。与ISBSG存储库的案例研究表明,所提出的调优过程结合一个简单的基于阈值的聚类方法是最成功的,使Dycom大幅减少(10倍)所需的WC训练项目的数量,同时保持(甚至提高)预测性能相比,相应的WC模型。提供了详细的分析,以了解这种方法是否有效的条件。总体而言,建议的在线监督调整程序是成功的,使一个非常简单的基于阈值的聚类方法,以获得最有竞争力的Dycom结果。这证明了以监督的方式随时间自动调整超参数的价值。
Software effort estimation is an online supervised learning problem, where new training projects may become available over time. In this scenario, the Cross-Company (CC) approach Dycom can drastically reduce the number of Within-Company (WC) projects needed for training, saving their collection cost. However, Dycom requires CC projects to be split into subsets. Both the number and composition of such subsets can affect Dycom’s predictive performance. Even though clustering methods could be used to automatically create CC subsets, there are no procedures for automatically tuning the number of clusters over time in online supervised scenarios. This paper proposes the first procedure for that. An investigation of Dycom using six clustering methods and three automated tuning procedures is performed, to check whether clustering with automated tuning can create well performing CC splits. A case study with the ISBSG Repository shows that the proposed tuning procedure in combination with a simple threshold-based clustering method is the most successful in enabling Dycom to drastically reduce (by a factor of 10) the number of required WC training projects, while maintaining (or even improving) predictive performance in comparison with a corresponding WC model. A detailed analysis is provided to understand the conditions under which this approach does or does not work well. Overall, the proposed online supervised tuning procedure was generally successful in enabling a very simple threshold-based clustering approach to obtain the most competitive Dycom results. This demonstrates the value of automatically tuning hyperparameters over time in a supervised way.