Blending Controllers via Multi-Objective Bandits
Blending Controllers via Multi-Objective Bandits
复制标题
DOI:
10.23919/acc53348.2022.9867486
复制
发表时间:
2020-07
期刊:
影响因子:
--
通讯作者:
Parham Gohari;Franck Djeumou;Abraham P. Vinod;U. Topcu
中科院分区:
文献类型:
--
作者:
Parham Gohari;Franck Djeumou;Abraham P. Vinod;U. Topcu
Performance and safety are often two competing objectives in decision-making problems. We study the problem of integrating a collection of controllers with different safety and performance levels into one that takes a middle-ground position amongst them. In the first contribution, we formulate the problem of blending controllers using the framework of constrained Markov decision processes and contextual multi-objective bandits. We use the reward function and the auxiliary costs of the Markov decision process to measure the performance and the safety of a controller, respectively. We subsequently use these measures to form the feedback of a bandit whose arms are the input controllers. The blending algorithm must interact with the bandit and minimize a regret term that measures the suboptimality of the pulled arms with respect to an expert whose choice of arms is Pareto optimal. In the second contribution, we design a blending algorithm and show that its average regret converges to zero. We also derive an upper bound on the algorithm’s suboptimality in performance and safety and we show that its computation imposes no additional computational complexity. We empirically demonstrate the algorithm’s success in blending a safe and a performant controller in a variety of Safety Gym environments. The results reflect the following key takeaway: the blended controller shows a strict improvement in performance compared to the safe controller and is safer than the performant controller.