Non-Stochastic Control with Bandit Feedback

Non-Stochastic Control with Bandit Feedback
复制标题

带强盗反馈的非随机控制

DOI:
--
复制
发表时间:
2020
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
Elad Hazan
Elad Hazan
中科院分区:
--
文献类型:
--
作者:
Paula Gradu;John Hallman;Elad Hazan

文献摘要

参考文献

被引文献

相似文献

我们研究了具有对抗扰动的线性动力系统的控制问题,其中对控制器可用的唯一反馈是标量损失,并且损失函数本身是未知的。对于这个问题,无论系统是已知的还是未知的,我们都给出了一个有效的次线性后悔算法。算法的主要困难在于损失对过去控制的依赖性。为了克服这个问题,我们提出了一个有效的算法的一般设置的强盗凸优化的损失函数的记忆,这可能是独立的利益。
We study the problem of controlling a linear dynamical system with adversarial perturbations where the only feedback available to the controller is the scalar loss, and the loss function itself is unknown. For this problem, with either a known or unknown system, we give an efficient sublinear regret algorithm. The main algorithmic difficulty is the dependence of the loss on past controls. To overcome this issue, we propose an efficient algorithm for the general setting of bandit convex optimization for loss functions with memory, which may be of independent interest.
DOI: --
发表时间: 2018-01
影响因子: 8.7
作者:
Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi
通讯作者: Maryam Fazel;Rong Ge;S. Kakade;M. Mesbahi