Quickly Generating Diverse Valid Test Inputs with Reinforcement Learning ICSE 2020

Quickly Generating Diverse Valid Test Inputs with Reinforcement Learning ICSE 2020
复制标题

使用强化学习快速生成多样化的有效测试输入 ICSE 2020

DOI:
10.1145/3380399
复制
发表时间:
2020
期刊:
International conference on software engineering (ICSE'2020
影响因子:
--
通讯作者:
Sameer Reddy, Caroline Lemieux
Sameer Reddy, Caroline Lemieux
中科院分区:
--
文献类型:
--
作者:
Sameer Reddy, Caroline Lemieux

文献摘要

相似文献

基于属性的测试是验证程序逻辑的流行方法。一个有效的基于属性的测试可以快速生成许多不同的有效测试输入,并通过参数化的测试驱动程序运行它们。然而,当测试驱动程序需要对输入进行严格的有效性约束时,完全随机的输入生成无法生成足够的有效输入。解决该问题的现有方法依赖于通过仪表化输入生成器和/或测试驱动器收集的白盒或灰盒信息。然而,收集这样的信息降低了可以执行测试的速度。在本文中,我们提出并研究了一个黑盒方法产生有效的测试输入。我们首先正式的问题,引导随机输入生成器产生一组不同的有效输入。这种形式化突出了guide的作用,guide控制着随机输入生成器中的选择空间。然后,我们提出了一个基于强化学习(RL)的解决方案,使用表格,基于策略的RL方法来指导生成器。我们评估这种方法,RLCheck,对纯随机输入生成,以及一个国家的最先进的灰箱进化算法,在四个现实世界的基准。我们发现,在相同的时间预算,RLCheck生成一个数量级更多样化的有效输入比基线。
Property-based testing is a popular approach for validating the logic of a program. An effective property-based test quickly generates many diverse valid test inputs and runs them through a parameterized test driver. However, when the test driver requires strict validity constraints on the inputs, completely random input generation fails to generate enough valid inputs. Existing approaches to solving this problem rely on whitebox or greybox information collected by instrumenting the input generator and/or test driver. However, collecting such information reduces the speed at which tests can be executed. In this paper, we propose and study a black-box approach for generating valid test inputs. We first formalize the problem of guiding random input generators towards producing a diverse set of valid inputs. This formalization highlights the role of aguidewhich governs the space of choices within a random input generator. We then propose a solution based on reinforcement learning (RL), using a tabular, on-policy RL approach to guide the generator. We evaluate this approach,RLCheck,against pure random input generation as well as a state-of-the-art greybox evolutionary algorithm, on four real-world benchmarks. We find that in the same time budget,RLCheckgenerates an order of magnitude more diverse valid inputs than the baselines.