Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures

Rigorous Agent Evaluation: An Adversarial Approach to Uncover Catastrophic Failures
复制标题

严格的代理评估:发现灾难性故障的对抗性方法

DOI:
--
复制
发表时间:
2018
期刊:
International Conference on Learning Representations
影响因子:
--
通讯作者:
Pushmeet Kohli
Pushmeet Kohli
中科院分区:
--
文献类型:
--
作者:
J. Uesato;Ananya Kumar;Csaba Szepesvari;Tom Erez;Avraham Ruderman;Keith Anderson;Krishnamurthy Dvijotham;N. Heess;Pushmeet Kohli

文献摘要

被引文献

相似文献

本文讨论了在自动驾驶等安全关键领域评估学习系统的问题,在这些领域,故障可能会产生灾难性的后果。我们专注于两个问题:搜索的情况下,学习代理失败,并评估其失败的概率。强化学习中的标准代理评估方法Vanilla Monte Carlo可能会完全忽略失败,导致部署不安全的代理。我们证明了这是当前代理的一个问题,即使匹配用于训练的计算有时也不足以进行评估。为了解决这个缺点,我们借鉴了罕见事件概率估计文献,并提出了一种对抗性评估方法。我们的方法侧重于评估不利选择的情况下,同时仍然提供无偏估计的故障概率。关键的困难在于识别这些对抗性的情况--因为失败是罕见的,所以几乎没有信号来驱动优化。为了解决这个问题,我们提出了一个连续的方法,学习相关的,但不太强大的代理故障模式。我们的方法还允许重复使用已收集的数据来训练代理。我们证明了对抗性评估在两个标准领域的有效性:仿人控制和模拟驾驶。实验结果表明,我们的方法可以找到灾难性的故障和估计代理的故障率比标准的评估方案快多个数量级,在几分钟到几小时,而不是几天。
This paper addresses the problem of evaluating learning systems in safety critical domains such as autonomous driving, where failures can have catastrophic consequences. We focus on two problems: searching for scenarios when learned agents fail and assessing their probability of failure. The standard method for agent evaluation in reinforcement learning, Vanilla Monte Carlo, can miss failures entirely, leading to the deployment of unsafe agents. We demonstrate this is an issue for current agents, where even matching the compute used for training is sometimes insufficient for evaluation. To address this shortcoming, we draw upon the rare event probability estimation literature and propose an adversarial evaluation approach. Our approach focuses evaluation on adversarially chosen situations, while still providing unbiased estimates of failure probabilities. The key difficulty is in identifying these adversarial situations -- since failures are rare there is little signal to drive optimization. To solve this we propose a continuation approach that learns failure modes in related but less robust agents. Our approach also allows reuse of data already collected for training the agent. We demonstrate the efficacy of adversarial evaluation on two standard domains: humanoid control and simulated driving. Experimental results show that our methods can find catastrophic failures and estimate failures rates of agents multiple orders of magnitude faster than standard evaluation schemes, in minutes to hours rather than days.