Reliability Assessment and Safety Arguments for Machine Learning Components in System Assurance

Reliability Assessment and Safety Arguments for Machine Learning Components in System Assurance
复制标题

DOI:
10.1145/3570918
复制
发表时间:
2021-11
影响因子:
2
通讯作者:
Yizhen Dong;Wei Huang;Vibhav Bharti;V. Cox;Alec Banks;Sen Wang;Xingyu Zhao;S. Schewe;Xiaowei Huang
Yizhen Dong;Wei Huang;Vibhav Bharti;V. Cox;Alec Banks;Sen Wang;Xingyu Zhao;S. Schewe;Xiaowei Huang
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yizhen Dong;Wei Huang;Vibhav Bharti;V. Cox;Alec Banks;Sen Wang;Xingyu Zhao;S. Schewe;Xiaowei Huang

文献摘要

相似文献

越来越多地使用嵌入在自主系统中的机器学习(ML)组件-所谓的学习支持系统(LESs)-导致迫切需要确保其功能安全性。至于传统的功能安全,工业界和学术界正在形成的共识是为此目的使用保证案例。通常,保证案例支持可靠性声明以支持安全性,并且可以被视为组织从安全分析和可靠性建模活动中生成的论点和证据的结构化方式。虽然这种保证活动传统上是由基于共识的标准指导的,这些标准是从丰富的工程经验中开发的,但由于ML模型的特性和设计,LES在安全关键应用中提出了新的挑战。在这篇文章中,我们首先提出了一个LES的整体保证框架,重点是定量方面,例如,将系统级安全目标分解为组件级要求,并支持可靠性度量中所述的声明。然后,我们介绍了一种新的模型无关的可靠性评估模型(RAM)的ML分类器,利用操作配置文件和鲁棒性验证证据。我们讨论了模型假设和评估我们的RAM所揭示的ML可靠性的固有挑战,并提出了实际使用的解决方案。基于RAM还开发了较低ML组件级别的概率安全参数模板。最后,为了评估和演示我们的方法,我们不仅在合成/基准数据集上进行实验,而且还将我们的方法与模拟自主水下航行器和物理无人地面车辆的案例研究相结合。
The increasing use of Machine Learning (ML) components embedded in autonomous systems—so-called Learning-Enabled Systems (LESs)—has resulted in the pressing need to assure their functional safety. As for traditional functional safety, the emerging consensus within both, industry and academia, is to use assurance cases for this purpose. Typically assurance cases support claims of reliability in support of safety, and can be viewed as a structured way of organising arguments and evidence generated from safety analysis and reliability modelling activities. While such assurance activities are traditionally guided by consensus-based standards developed from vast engineering experience, LESs pose new challenges in safety-critical application due to the characteristics and design of ML models. In this article, we first present an overall assurance framework for LESs with an emphasis on quantitative aspects, e.g., breaking down system-level safety targets to component-level requirements and supporting claims stated in reliability metrics. We then introduce a novel model-agnostic Reliability Assessment Model (RAM) for ML classifiers that utilises the operational profile and robustness verification evidence. We discuss the model assumptions and the inherent challenges of assessing ML reliability uncovered by our RAM and propose solutions to practical use. Probabilistic safety argument templates at the lower ML component-level are also developed based on the RAM. Finally, to evaluate and demonstrate our methods, we not only conduct experiments on synthetic/benchmark datasets but also scope our methods with case studies on simulated Autonomous Underwater Vehicles and physical Unmanned Ground Vehicles.