Deep Reinforcement Learning Verification: A Survey

Deep Reinforcement Learning Verification: A Survey
复制标题

DOI:
10.1145/3596444
复制
发表时间:
2023-05
影响因子:
16.6
通讯作者:
Matthew Landers;Afsaneh Doryab
Matthew Landers;Afsaneh Doryab
中科院分区:
计算机科学1区
文献类型:
--
作者:
Matthew Landers;Afsaneh Doryab

文献摘要

相似文献

深度强化学习(DRL)已被证明能够在许多复杂任务上表现出超人的性能。为了实现这一成功,DRL算法训练决策代理来选择最大化一些长期性能指标的动作。然而,在现实世界的许多领域中,最佳性能并不足以证明算法的使用-例如,有时必须严格确保系统的鲁棒性,稳定性或安全性。因此,已经出现了用于验证DRL系统的方法。这些算法可以保证系统在无限输入集上的属性,但任务并不简单。DRL依赖于深度神经网络(DNN)。DNN通常被称为“黑匣子”,因为检查它们各自的结构并不能阐明它们的决策过程。此外,DRL用于解决的问题的顺序性质促进了显著的可扩展性挑战。最后,由于DRL环境通常是随机的,验证方法必须考虑概率行为。为了解决这些问题,出现了一个新的子领域。在这次调查中,我们建立了DRL和DRL验证的基础,定义了DRL验证方法的分类,描述了处理随机性的方法,描述了与编写规范相关的考虑因素,列举了常见的测试任务/环境,并详细介绍了未来研究的机会。
Deep reinforcement learning (DRL) has proven capable of superhuman performance on many complex tasks. To achieve this success, DRL algorithms train a decision-making agent to select the actions that maximize some long-term performance measure. In many consequential real-world domains, however, optimal performance is not enough to justify an algorithm’s use—for example, sometimes a system’s robustness, stability, or safety must be rigorously ensured. Thus, methods for verifying DRL systems have emerged. These algorithms can guarantee a system’s properties over an infinite set of inputs, but the task is not trivial. DRL relies on deep neural networks (DNNs). DNNs are often referred to as “black boxes” because examining their respective structures does not elucidate their decision-making processes. Moreover, the sequential nature of the problems DRL is used to solve promotes significant scalability challenges. Finally, because DRL environments are often stochastic, verification methods must account for probabilistic behavior. To address these complications, a new subfield has emerged. In this survey, we establish the foundations of DRL and DRL verification, define a taxonomy for DRL verification methods, describe approaches for dealing with stochasticity, characterize considerations related to writing specifications, enumerate common testing tasks/environments, and detail opportunities for future research.