Non-Stochastic Control with Bandit Feedback
Non-Stochastic Control with Bandit Feedback
复制标题
带强盗反馈的非随机控制
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Elad Hazan
中科院分区:
文献类型:
--
作者:
Paula Gradu;John Hallman;Elad Hazan
We study the problem of controlling a linear dynamical system with adversarial perturbations where the only feedback available to the controller is the scalar loss, and the loss function itself is unknown. For this problem, with either a known or unknown system, we give an efficient sublinear regret algorithm. The main algorithmic difficulty is the dependence of the loss on past controls. To overcome this issue, we propose an efficient algorithm for the general setting of bandit convex optimization for loss functions with memory, which may be of independent interest.
影响因子:
8.7
作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi
通讯作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi