Decision Making with STPA through Markov Decision Process, a Theoretic Framework for Safe Human-Robot Collaboration

Decision Making with STPA through Markov Decision Process, a Theoretic Framework for Safe Human-Robot Collaboration
复制标题

通过马尔可夫决策过程(安全人机协作的理论框架)进行 STPA 决策

DOI:
10.3390/app11115212
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
I. Dokas
I. Dokas
中科院分区:
--
文献类型:
--
作者:
Angeliki Zacharaki;I. Kostavelis;I. Dokas

文献摘要

被引文献

相似文献

在过去的几十年中,能够在笼子外操作的协作机器人被广泛用于工业中,以帮助人类完成平凡而苛刻的制造任务。虽然这种机器人在设计上具有固有的安全性,但它们通常配有外部传感器和其他网络物理系统,以促进与人类的密切合作,这往往使协作生态系统不安全,容易发生危险。我们介绍了一种方法,利用部分可观察马尔可夫决策过程(POMDP)合并系统的名义行动沿着与不安全的控制行动所造成的系统理论过程分析(STPA)。不断提示系统进入更安全状态的决策机制通过将系统安全意识与特定的选定动作组相关联来提供关于协作生态系统的安全级别的情况意识来实现。POMDP补偿了协作环境当前状态的部分可观测性和不确定性,并创建了安全筛选策略,该策略倾向于在操作阶段期间在真实的时间内做出使系统从不安全状态平衡到安全状态的决策。的理论框架进行了评估,在一个模拟的人机协作的情况下,并证明能够识别失败和成功的情况。
During the last decades, collaborative robots capable of operating out of their cages are widely used in industry to assist humans in mundane and harsh manufacturing tasks. Although such robots are inherently safe by design, they are commonly accompanied by external sensors and other cyber-physical systems, to facilitate close cooperation with humans, which frequently render the collaborative ecosystem unsafe and prone to hazards. We introduce a method that capitalizes on partially observable Markov decision processes (POMDP) to amalgamate nominal actions of the system along with unsafe control actions posed by the System Theoretic Process Analysis (STPA). A decision-making mechanism that constantly prompts the system into a safer state is realized by providing situation awareness about the safety levels of the collaborative ecosystem by associating the system safety awareness with specific groups of selected actions. POMDP compensates the partial observability and uncertainty of the current state of the collaborative environment and creates safety screening policies that tend to make decisions that balance the system from unsafe to safe states in real time during the operational phase. The theoretical framework is assessed on a simulated human–robot collaborative scenario and proved capable of identifying loss and success scenarios.