Sparse reward for reinforcement learning-based continuous integration testing

Sparse reward for reinforcement learning-based continuous integration testing
复制标题

基于强化学习的持续集成测试的稀疏奖励

DOI:
10.1002/smr.2409
复制
发表时间:
2021
期刊:
Journal of Software: Evolution and Process
影响因子:
--
通讯作者:
Li Qianyu
Li Qianyu
中科院分区:
其他
文献类型:
--
作者:
Yang Yang;Li Zheng;Shang Ying;Li Qianyu

文献摘要

相似文献

强化学习(RL)已被用于优化持续集成(CI)测试,其中奖励在指导测试用例优先级(TCP)策略的调整中起着关键作用。在CI测试中,集成的频率通常很高,而测试用例的失败率很低。因此,强化学习在CI测试中会获得较少的奖励,这可能导致强化学习的学习效率较低,甚至难以收敛。本文引入了三种奖励来解决RL在CI测试中的稀疏奖励问题。首先,定义基于历史失败密度的奖励(HFD),它客观地表示了稀疏奖励问题。其次,提出了基于平均失败位置的奖励(AFP),以增加奖励值,减少稀疏奖励的影响。在此基础上,提出了一种基于附加奖励的技术,提取通过测试用例的测试发生频率作为附加奖励。对14个真实行业数据集进行了实证研究。实验结果令人鼓舞,特别是附加奖励的奖励可以提高NAPFD(归一化平均故障检测百分比)高达21.97%,提高Recall(召回率)高达21.87%,平均提高TTF(测试失败)9.99个位置。
Reinforcement learning (RL) has been used to optimize the continuous integration (CI) testing, where the reward plays a key role in directing the adjustment of the test case prioritization (TCP) strategy. In CI testing, the frequency of integration is usually very high, while the failure rate of test cases is low. Consequently, RL will get scarce rewards in CI testing, which may lead to low learning efficiency of RL and even difficulty in convergence. This paper introduces three rewards to tackle the issue of sparse rewards of RL in CI testing. First, the historical failure density‐based reward (HFD) is defined, which objectively represents the sparse reward problem. Second, the average failure position‐based reward (AFP) is proposed to increase the reward value and reduce the impact of sparse rewards. Furthermore, a technique based on additional reward is proposed, which extracts the test occurrence frequency of passed test cases for additional rewards. Empirical studies are conducted on 14 real industry data sets. The experiment results are promising, especially the reward with additional reward can improve NAPFD (Normalized Average Percentage of Faults Detected) by up to 21.97%, enhance Recall with a maximum of 21.87%, and increase TTF (Test to Fail) by an average of 9.99 positions.