Certifiable Evaluation for Autonomous Vehicle Perception Systems using Deep Importance Sampling (Deep IS)

Certifiable Evaluation for Autonomous Vehicle Perception Systems using Deep Importance Sampling (Deep IS)
复制标题

DOI:
10.1109/itsc55140.2022.9922202
复制
发表时间:
2022-10
期刊:
2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC)
影响因子:
--
通讯作者:
Mansur Arief;Zhepeng Cen;Zhen-Yan Liu;Zhiyuan Huang;Bo Li;H. Lam;Ding Zhao
Mansur Arief;Zhepeng Cen;Zhen-Yan Liu;Zhiyuan Huang;Bo Li;H. Lam;Ding Zhao
中科院分区:
其他
文献类型:
--
作者:
Mansur Arief;Zhepeng Cen;Zhen-Yan Liu;Zhiyuan Huang;Bo Li;H. Lam;Ding Zhao

文献摘要

相似文献

在自然条件下高精度地评估自动驾驶汽车(AV)的性能及其复杂的人工智能驱动功能仍然是一个挑战,特别是在故障或危险情况很少的情况下。稀有性不仅需要巨大的样本量才能实现高置信度的剩余风险估计,而且还可能导致难以检测的严重风险低估问题。同时,最先进的罕见的安全关键事件评估方法,具有正确性保证,可以计算在一定条件下的真实风险的上限,这限制了它的实际应用。在这项工作中,我们提出了深度重要性采样(Deep IS)框架,该框架利用深度神经网络来获得有效的偏差较小的风险估计,其效率与最先进的方法相当。在评估交通标志分类器的错误分类率的数值实验中,Deep IS仅需要朴素采样方法所需样本的1/40即可实现10%的相对误差。此外,与风险上限相比,Deep IS产生的估计值保守性低10倍,与真实目标的差异最多仅为10%。这种高效的基于深度学习的IS程序有望成为一种高效的方法,用于处理通常高维的功能安全问题,这些问题在AV领域中普遍存在罕见的自然故障情况。
Evaluating the performance of autonomous vehicles (AV) and their complex AI-driven functionalities to high precision under naturalistic conditions remains a challenge, especially when the failure or dangerous cases are rare. Rarity does not only require an enormous sample size for a naive method to achieve high confidence residual risk estimation, but it can also cause serious risk underestimation issues that is hard to detect. Meanwhile, the state-of-the-art rare safety-critical event evaluation approach that comes with a correctness guarantee can compute an upper bound for the true risk under certain conditions, which limits its practical uses. In this work, we propose Deep Importance Sampling (Deep IS) framework that utilizes a deep neural network to obtain an efficient less biased risk estimate, with an efficiency that is on par with that of the state-of-the-art method. In the numerical experiment evaluating the misclassification rate of a traffic sign classifier, Deep IS only needs 1/40-th of the samples required by a naive sampling method to achieve 10% relative error. Furthermore, the estimate produced by Deep IS is 10 times less conservative compared to the risk upper bound and only off by at most 10% difference to the true target. This efficient deep-learning-based IS procedure promises a highly efficient method to deal with often high-dimensional functional safety problems with rare naturalistic failure cases that are prevalent in AV domains.