Index-aware reinforcement learning for adaptive video streaming at the wireless edge
Index-aware reinforcement learning for adaptive video streaming at the wireless edge
复制标题
DOI:
10.1145/3492866.3549726
复制
发表时间:
2022-10
期刊:
影响因子:
--
通讯作者:
Guojun Xiong;Xudong Qin;B. Li;Rahul Singh;J. Li
中科院分区:
文献类型:
--
作者:
Guojun Xiong;Xudong Qin;B. Li;Rahul Singh;J. Li
We study adaptive video streaming for multiple users in wireless access edge networks with unreliable channels. The key challenge is to jointly optimize the video bitrate adaptation and resource allocation such that the users' cumulative quality of experience is maximized. This problem is a finite-horizon restless multi-armed multi-action bandit problem and is provably hard to solve. To overcome this challenge, we propose a computationally appealing index policy entitled Quality Index Policy, which is well-defined without the Whittle indexability condition and is provably asymptotically optimal without the global attractor condition. These two conditions are widely needed in the design of most existing index policies, which are difficult to establish in general. Since the wireless access edge network environment is highly dynamic with system parameters unknown and time-varying, we further develop an index-aware reinforcement learning (RL) algorithm dubbed QA-UCB. We show that QA-UCB achieves a sub-linear regret with a low-complexity since it fully exploits the structure of the Quality Index Policy for making decisions. Extensive simulations using real-world traces demonstrate significant gains of proposed policies over conventional approaches. We note that the proposed framework for designing index policy and index-aware RL algorithm is of independent interest and could be useful for other large-scale multi-user problems.