Learning Adversarial Markov Decision Processes with Bandit Feedback and Unknown Transition

Learning Adversarial Markov Decision Processes with Bandit Feedback and Unknown Transition
复制标题

DOI:
--
复制
发表时间:
2019-12
期刊:
--
影响因子:
--
通讯作者:
Chi Jin;Tiancheng Jin;Haipeng Luo;S. Sra;Tiancheng Yu
Chi Jin;Tiancheng Jin;Haipeng Luo;S. Sra;Tiancheng Yu
中科院分区:
其他
文献类型:
--
作者:
Chi Jin;Tiancheng Jin;Haipeng Luo;S. Sra;Tiancheng Yu

文献摘要

被引文献

相似文献

研究了转移函数未知、具有强盗反馈和对手损失的情景有限时间马氏决策过程的学习问题。我们提出了一个高效的算法,它以很高的概率实现了$\数学上的{\tilde{O}}(L|X|^2\SQRT{|A|T})$遗憾,其中$L$是地平线,$|X|$是状态数,$|A|$是动作数,$T$是情节数。就我们所知,我们的算法是第一个在这种具有挑战性的环境中确保{$\mathcal{\tilde{O}}(\sqrt{T})$}后悔的算法。我们的主要技术贡献是引入了一个乐观的损失估计器,该估计器由$\textit{上界}$反向加权。
We consider the problem of learning in episodic finite-horizon Markov decision processes with unknown transition function, bandit feedback, and adversarial losses. We propose an efficient algorithm that achieves $\mathcal{\tilde{O}}(L|X|^2\sqrt{|A|T})$ regret with high probability, where $L$ is the horizon, $|X|$ is the number of states, $|A|$ is the number of actions, and $T$ is the number of episodes. To the best of our knowledge, our algorithm is the first one to ensure {$\mathcal{\tilde{O}}(\sqrt{T})$} regret in this challenging setting. Our key technical contribution is to introduce an optimistic loss estimator that is inversely weighted by an $\textit{upper occupancy bound}$.