A Near-Optimal Best-of-Both-Worlds Algorithm for Online Learning with Feedback Graphs

A Near-Optimal Best-of-Both-Worlds Algorithm for Online Learning with Feedback Graphs
复制标题

一种近乎最优的两全其美的反馈图在线学习算法

DOI:
10.48550/arxiv.2206.00557
复制
发表时间:
2022
期刊:
ArXiv
影响因子:
--
通讯作者:
Yevgeny Seldin
Yevgeny Seldin
中科院分区:
--
文献类型:
--
作者:
Chloé Rouyer;Dirk van der Hoeven;Nicolò Cesa;Yevgeny Seldin

文献摘要

参考文献

被引文献

相似文献

我们考虑使用反馈图的在线学习,这是一个顺序决策框架,其中学习者的反馈由动作集上的有向图决定。我们提出了一种计算效率高的算法,用于在该框架中学习,同时在随机和对抗环境中实现近最优后悔界限。与遗忘对手的界限是$\tilde{O} (\sqrt{\alpha T})$,其中$T$是时间范围,$\alpha$是反馈图的独立性数。随机环境的界是$O\big( (\ln T)^2 \max_{S\in \mathcal I(G)} \sum_{i \in S} \Delta_i^{-1}\big)$,其中$\mathcal I(G)$是图的适当定义的无向版本中所有独立集的族,$\Delta_i$是次优性间隙。该算法结合了EXP3++算法和EXP3的思想。反馈图的G算法与一种新的探索方案。该方案利用图的结构来减少探索,是获得反馈图两全其美保证的关键。我们还将算法和结果扩展到允许反馈图随时间变化的设置。
We consider online learning with feedback graphs, a sequential decision-making framework where the learner's feedback is determined by a directed graph over the action set. We present a computationally efficient algorithm for learning in this framework that simultaneously achieves near-optimal regret bounds in both stochastic and adversarial environments. The bound against oblivious adversaries is $\tilde{O} (\sqrt{\alpha T})$, where $T$ is the time horizon and $\alpha$ is the independence number of the feedback graph. The bound against stochastic environments is $O\big( (\ln T)^2 \max_{S\in \mathcal I(G)} \sum_{i \in S} \Delta_i^{-1}\big)$ where $\mathcal I(G)$ is the family of all independent sets in a suitably defined undirected version of the graph and $\Delta_i$ are the suboptimality gaps. The algorithm combines ideas from the EXP3++ algorithm for stochastic and adversarial bandits and the EXP3.G algorithm for feedback graphs with a novel exploration scheme. The scheme, which exploits the structure of the graph to reduce exploration, is key to obtain best-of-both-worlds guarantees with feedback graphs. We also extend our algorithm and results to a setting where the feedback graphs are allowed to change over time.
通过根对数正则化器的最小最大最优分位数和半对抗性遗憾
DOI: --
发表时间: 2021
期刊: Advances in neural information processing systems
影响因子: --
作者:
Negrea, Jeffrey;Bilodeau, Blair;Campolongo, Nicolò;Orabona, Francesco;Roy, Dan
通讯作者: Roy, Dan