Blending Controllers via Multi-Objective Bandits

Blending Controllers via Multi-Objective Bandits
复制标题

DOI:
10.23919/acc53348.2022.9867486
复制
发表时间:
2020-07
期刊:
2022 American Control Conference (ACC)
影响因子:
--
通讯作者:
Parham Gohari;Franck Djeumou;Abraham P. Vinod;U. Topcu
Parham Gohari;Franck Djeumou;Abraham P. Vinod;U. Topcu
中科院分区:
其他
文献类型:
--
作者:
Parham Gohari;Franck Djeumou;Abraham P. Vinod;U. Topcu

文献摘要

相似文献

在决策问题中,性能和安全往往是两个相互竞争的目标。我们研究的问题,集成到一个需要一个中间地带的位置,其中一个不同的安全和性能水平的控制器的集合。在第一个贡献中,我们制定了混合控制器使用的框架约束马尔可夫决策过程和上下文多目标土匪的问题。我们使用马尔可夫决策过程的报酬函数和辅助成本来衡量控制器的性能和安全性。随后,我们使用这些措施来形成一个强盗的反馈,其手臂是输入控制器。混合算法必须与强盗互动,并最小化一个后悔项,该后悔项衡量被拉手臂相对于其选择手臂是帕累托最优的专家的次优性。在第二个贡献中,我们设计了一个混合算法,并证明了它的平均遗憾收敛到零。我们还推导出一个上界的算法的次优性能和安全性,我们表明,它的计算不施加额外的计算复杂性。我们经验证明了该算法的成功融合安全和性能控制器在各种安全健身房环境。结果反映了以下关键要点:与安全控制器相比,混合控制器在性能上有严格的改进,并且比性能控制器更安全。
Performance and safety are often two competing objectives in decision-making problems. We study the problem of integrating a collection of controllers with different safety and performance levels into one that takes a middle-ground position amongst them. In the first contribution, we formulate the problem of blending controllers using the framework of constrained Markov decision processes and contextual multi-objective bandits. We use the reward function and the auxiliary costs of the Markov decision process to measure the performance and the safety of a controller, respectively. We subsequently use these measures to form the feedback of a bandit whose arms are the input controllers. The blending algorithm must interact with the bandit and minimize a regret term that measures the suboptimality of the pulled arms with respect to an expert whose choice of arms is Pareto optimal. In the second contribution, we design a blending algorithm and show that its average regret converges to zero. We also derive an upper bound on the algorithm’s suboptimality in performance and safety and we show that its computation imposes no additional computational complexity. We empirically demonstrate the algorithm’s success in blending a safe and a performant controller in a variety of Safety Gym environments. The results reflect the following key takeaway: the blended controller shows a strict improvement in performance compared to the safe controller and is safer than the performant controller.