A Novel Approach to Error Resilience in Online Reinforcement Learning

A Novel Approach to Error Resilience in Online Reinforcement Learning
复制标题

DOI:
10.1109/iolts59296.2023.10224892
复制
发表时间:
2023-07
期刊:
2023 IEEE 29th International Symposium on On-Line Testing and Robust System Design (IOLTS)
影响因子:
--
通讯作者:
C. Amarnath;A. Chatterjee
C. Amarnath;A. Chatterjee
中科院分区:
其他
文献类型:
--
作者:
C. Amarnath;A. Chatterjee

文献摘要

相似文献

基于在线强化学习 (RL) 的系统越来越多地部署在从无人机控制到医疗机器人等各种安全关键应用中。这些系统通常使用机载强化学习,而不是依赖高性能数据中心的远程操作。由于工作环境的动态特性,板载强化学习硬件很容易受到辐射、热效应和电噪声等软错误的影响,从而破坏计算结果。机器学习系统中现有的在线错误恢复方法依赖于大型训练数据集的可用性来配置恢复参数,这对于在线强化学习系统不一定可行。同样,涉及专用硬件或训练算法修改的其他方法也很难在机载强化学习应用中实现。相比之下,我们提出了一种新颖的在线强化学习错误恢复方法,该方法利用在(实时)强化学习训练过程中收集的运行统计数据来配置错误检测阈值,而无需访问参考训练数据集。在这种方法中,利用运行统计数据的统计浓度界限用于诊断神经元输出是否错误。然后这些错误的神经元被设置为零(被抑制)。我们的方法与最先进的方法进行了比较,并在几种 RL 算法上进行了验证,这些算法涉及在 CPU 和 GPU 硬件上使用多个浓度范围。
Online reinforcement learning (RL) based systems are being increasingly deployed in a variety of safety-critical applications ranging from drone control to medical robotics. These systems typically use RL onboard rather than relying on remote operation from high-performance datacenters. Due to the dynamic nature of the environments they work in, onboard RL hardware is vulnerable to soft errors from radiation, thermal effects and electrical noise that corrupt the results of computations. Existing approaches to on-line error resilience in machine learning systems have relied on availability of the large training datasets to configure resilience parameters, which is not necessarily feasible for online RL systems. Similarly, other approaches involving specialized hardware or modifications to training algorithms are difficult to implement for onboard RL applications. In contrast, we present a novel error resilience approach for online RL that makes use of running statistics collected across the (real-time) RL training process to configure error detection thresholds without the need to access a reference training dataset. In this methodology, statistical concentration bounds leveraging running statistics are used to diagnose neuron outputs as erroneous. These erroneous neurons are then set to zero (suppressed). Our approach is compared against the state of the art and validated on several RL algorithms involving the use of multiple concentration bounds on CPU as well as GPU hardware.