An Industrial Evaluation of Unit Test Generation: Finding Real Faults in a Financial Application

An Industrial Evaluation of Unit Test Generation: Finding Real Faults in a Financial Application
复制标题

DOI:
10.1109/icse-seip.2017.27
复制
发表时间:
2017-05
期刊:
2017 IEEE/ACM 39th International Conference on Software Engineering: Software Engineering in Practice Track (ICSE-SEIP)
影响因子:
--
通讯作者:
M. Almasi;H. Hemmati;G. Fraser;Andrea Arcuri;Janis Benefelds
M. Almasi;H. Hemmati;G. Fraser;Andrea Arcuri;Janis Benefelds
中科院分区:
其他
文献类型:
--
作者:
M. Almasi;H. Hemmati;G. Fraser;Andrea Arcuri;Janis Benefelds

文献摘要

被引文献

相似文献

近年来,在文献中已经对自动化单元测试的生成进行了广泛的研究。先前对开源系统的研究表明,测试生成工具在检测故障方面非常有效,但是它们在工业应用中的有效性和适用性?在本文中,我们使用SEB Life&Pension Handing Ab Riga分支拥有的人寿保险和养老金产品计算器引擎进行了调查。为了研究故障发现效果,我们从该软件项目的版本历史中提取了25个真正的故障,并为Java,EvoSuite和Randoop应用了两个最新的单元测试生成工具,它们实现了基于搜索的和反馈指导的随机测试产生分别。这些故障的自动生成的测试套件最多可检测到56.40%(evosuite)和38.00%(randoop)。对我们结果的分析表明,为了改善测试生成工具中的故障检测,需要解决的挑战。特别是,未检测到的故障的分类表明,其中97.62%取决于“特定原始值”(50.00%)或构造“对象的复杂状态配置”(47.62%)。为了研究适用性,我们调查了正在测试其经验和对测试生成工具和生成的测试案例的意见测试的开发人员。这导致了对从学术研究成功转移到工业实践的成功技术转移的学术原型的见解,例如需要与流行的构建工具集成,并提高生成的测试的可读性。
Automated unit test generation has been extensively studied in the literature in recent years. Previous studies on open source systems have shown that test generation tools are quite effective at detecting faults, but how effective and applicable are they in an industrial application? In this paper, we investigate this question using a life insurance and pension products calculator engine owned by SEB Life & Pension Holding AB Riga Branch. To study fault-finding effectiveness, we extracted 25 real faults from the version history of this software project, and applied two up-to-date unit test generation tools for Java, EVOSUITE and RANDOOP, which implement search-based and feedback-directed random test generation, respectively. Automatically generated test suites detected up to 56.40% (EVOSUITE) and 38.00% (RANDOOP) of these faults. The analysis of our results demonstrates challenges that need to be addressed in order to improve fault detection in test generation tools. In particular, classification of the undetected faults shows that 97.62% of them depend on either "specific primitive values" (50.00%) or the construction of "complex state configuration of objects" (47.62%). To study applicability, we surveyed the developers of the application under test on their experience and opinions about the test generation tools and the generated test cases. This leads to insights on requirements for academic prototypes for successful technology transfer from academic research to industrial practice, such as a need to integrate with popular build tools, and to improve the readability of the generated tests.