Engineering resilience into safety-critical systems

Engineering resilience into safety-critical systems
复制标题

将弹性工程融入安全关键系统

DOI:
--
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
B. Barrett
B. Barrett
中科院分区:
--
文献类型:
--
作者:
N. Leveson;Nicolas Dulac;D. Zipkin;J. Cutcher;John Carroll;B. Barrett

文献摘要

被引文献

相似文献

1 复原力和安全性 复原力通常是指在发生重大事故或事件后继续运营或恢复稳定状态的能力。这个定义侧重于复原力的反应性质和在沮丧后恢复的能力。在本章中,我们使用一个更通用的定义,其中包括预防不安。在我们的概念中,弹性是系统预防或适应不断变化的条件以维持(控制)系统属性的能力。在本章中,我们关心的财产是安全或风险。为了确保安全,系统必须具有弹性,可以避免故障和损失,并在事后做出适当的响应。在重大事故发生之前,组织通常会陷入风险增加的状态,直到导致损失的事件发生[12]。我们的目标是确定如何设计弹性系统,以应对导致向更高风险状态漂移的压力和影响,或者,如果不可能,则设计持续的风险管理系统来检测漂移并协助在损失事件发生之前制定适当的应对措施。我们的方法依赖于对社会技术系统进行建模和分析,并使用在设计社会技术系统时获得的信息,评估对事件的计划响应和建议的组织政策以防止不利的组织漂移,以及定义适当的指标来检测风险的变化(相当于“煤矿里的金丝雀”)。为了发挥作用,此类建模和分析必须能够处理具有分布式人力和自动化控制的复杂、紧密耦合的系统、先进技术和软件密集型系统以及系统的组织和社会方面。为此,我们使用基于系统论的新事故因果关系模型 (STAMP)。 STAMP 包括非线性、间接和反馈关系,与传统的因果关系和事故模型相比,可以更好地处理当今系统的复杂性和技术创新水平。在下一节中,我们将简要介绍 STAMP。然后,我们展示如何使用 STAMP 模型来设计和分析弹性,并将其应用于 NASA 航天飞机计划的安全文化。 *本章描述的研究部分得到了 NASA/USRA 计划/项目管理研究中心的资助。
1 Resilience and Safety Resilience is often defined in terms of the ability to continue operations or recover a stable state after a major mishap or event. This definition focuses on the reactive nature of resilience and the ability to recover after an upset. In this chaper, we use a more general definition that includes prevention of upsets. In our conception, resilience is the ability of systems to prevent or adapt to changing conditions in order to maintain (control over) a system property. In this chapter, the property we are concerned about is safety or risk. To ensure safety, the system must be resilient in terms of avoiding failures and losses, as well as responding appropriately after the fact. Major accidents are usually preceded by periods where the organization drifts toward states of increasing risk until the events occur that lead to a loss [12]. Our goal is to determine how to design resilient systems that respond to the pressures and influences causing the drift to states of higher risk or, if that is not possible, to design continuous risk management systems to detect the drift and assist in formulating appropriate responses before the loss event occurs. Our approach rests on modeling and analyzing socio-technical systems and using the information gained in designing the socio-technical system, in evaluating both planned responses to events and suggested organizational policies to prevent adverse organizational drift, and in defining appropriate metrics to detect changes in risk (the equivalent of a “canary in the coal mine”). To be useful, such modeling and analysis must be able to handle complex, tightly coupled systems with distributed human and automated control, advanced technology and software-intensive systems, and the organizational and social aspects of systems. To do this, we use a new model of accident causation (STAMP) based on system theory. STAMP includes non-linear, indirect, and feedback relationships and can better handle the levels of complexity and technological innovation in today’s systems than traditional causality and accident models. In the next section, we briefly describe STAMP. Then we show how STAMP models can be used to design and analyze resilience by applying it to the safety culture of the NASA Space Shuttle program. ∗The research described in this chapter was partially supported by a grant from the NASA/USRA Center for Program/Project Management Research.