Search-based software library recommendation using multi-objective optimization

Search-based software library recommendation using multi-objective optimization
复制标题

DOI:
10.1016/j.infsof.2016.11.007
复制
发表时间:
2017-03
期刊:
Inf. Softw. Technol.
影响因子:
--
通讯作者:
Ali Ouni;R. Kula;Marouane Kessentini;T. Ishio;D. Germán;Katsuro Inoue
Ali Ouni;R. Kula;Marouane Kessentini;T. Ishio;D. Germán;Katsuro Inoue
中科院分区:
其他
文献类型:
--
作者:
Ali Ouni;R. Kula;Marouane Kessentini;T. Ishio;D. Germán;Katsuro Inoue

文献摘要

被引文献

相似文献

内容:软件库重用显著提高了软件开发人员的生产力,缩短了上市时间,提高了软件质量和可重用性。然而,随着代码库中可重用软件库数量的不断增加,寻找和采用相关的软件库成为一个挑剔和复杂的task for developer.Objective:在本文中,我们提出了一种新的方法,称为LibFinder,以防止错过重用的机会,在软件维护和演化。我们的目标是为开发人员提供决策支持,以便轻松找到“有用的”第三方库来实现他们的软件系统。方法:为此,我们使用了非支配排序遗传算法(NSGA-II),一种基于多目标搜索的算法,以找到三个目标之间的权衡:1)最大化候选库与给定系统所使用的实际库之间的共同使用,2)最大化候选库与系统的源代码之间的语义相似性,结果:我们对来自Maven Central超级存储库的6083个不同的库进行了评估,这些库被从Github超级存储库获得的32,760个客户端系统使用。我们的研究结果表明,我们的方法优于其他三个现有的搜索技术和最先进的方法,而不是基于启发式搜索,并成功地推荐有用的图书馆在92%的准确率得分,51%的精度和召回率为68%,同时找到最佳的权衡考虑的三个目标。此外,我们通过对开发人员的两个工业Java系统的实证研究来评估我们的方法在实践中的有用性。结果显示,最初的开发者对推荐的前10个库的平均评分为3.25分(满分为5分)。结论:本研究建议:(1)从不同客户端系统收集的库使用历史和(2)库标识符中体现的库语义/内容应该平衡在一起,以实现有效的库推荐技术。
Context: Software library reuse has significantly increased the productivity of software developers, reduced time-to-market and improved software quality and reusability. However, with the growing number of reusable software libraries in code repositories, finding and adopting a relevant software library becomes a fastidious and complex task for developers.Objective: In this paper, we propose a novel approach calledLibFinderto prevent missed reuse opportunities during software maintenance and evolution. The goal is to provide a decision support for developers to easily find “useful” third-party libraries to the implementation of their software systems.Method: To this end, we used the non-dominated sorting genetic algorithm (NSGA-II), a multi-objective search-based algorithm, to find a trade-off between three objectives : 1) maximizing co-usage between a candidate library and the actual libraries used by a given system, 2) maximizing the semantic similarity between a candidate library and the source code of the system, and 3) minimizing the number of recommended libraries.Results: We evaluated our approach on 6083 different libraries from Maven Central super repository that were used by 32,760 client systems obtained from Github super repository. Our results show that our approach outperforms three other existing search techniques and a state-of-the art approach, not based on heuristic search, and succeeds in recommending useful libraries at an accuracy score of 92%, precision of 51% and recall of 68%, while finding the best trade-off between the three considered objectives. Furthermore, we evaluate the usefulness of our approach in practice through an empirical study on two industrial Java systems with developers. Results show that the top 10 recommended libraries was rated by the original developers with an average of 3.25 out of 5.Conclusion: This study suggests that (1) library usage history collected from different client systems and (2) library semantics/content embodied in library identifiers should be balanced together for an efficient library recommendation technique.