Index-aware reinforcement learning for adaptive video streaming at the wireless edge

Index-aware reinforcement learning for adaptive video streaming at the wireless edge
复制标题

DOI:
10.1145/3492866.3549726
复制
发表时间:
2022-10
期刊:
Proceedings of the Twenty-Third International Symposium on Theory, Algorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing
影响因子:
--
通讯作者:
Guojun Xiong;Xudong Qin;B. Li;Rahul Singh;J. Li
Guojun Xiong;Xudong Qin;B. Li;Rahul Singh;J. Li
中科院分区:
其他
文献类型:
--
作者:
Guojun Xiong;Xudong Qin;B. Li;Rahul Singh;J. Li

文献摘要

相似文献

我们研究了无线接入边缘网络中的多用户自适应视频流与不可靠信道。关键的挑战是联合优化视频比特率适配和资源分配,使得用户的累积体验质量最大化。这个问题是一个有限时域的不安分多臂多行动的强盗问题,是很难解决的。为了克服这一挑战,我们提出了一个计算上有吸引力的指数政策,称为质量指数政策,这是定义良好的没有惠特尔指数条件,并证明是渐近最优的没有全局吸引子条件。这两个条件在大多数现有指数政策的设计中是广泛需要的,一般难以建立。由于无线接入边缘网络环境是高度动态的,系统参数未知且时变,我们进一步开发了一种索引感知强化学习(RL)算法,称为QA-UCB。我们表明,QA-UCB实现了一个低复杂度的次线性遗憾,因为它充分利用了结构的质量指标政策的决策。使用真实世界的痕迹进行了广泛的模拟,表明所提出的政策比传统方法有显着的收益。我们注意到,所提出的框架设计索引策略和索引感知RL算法是独立的利益,并可能是有用的其他大规模的多用户问题。
We study adaptive video streaming for multiple users in wireless access edge networks with unreliable channels. The key challenge is to jointly optimize the video bitrate adaptation and resource allocation such that the users' cumulative quality of experience is maximized. This problem is a finite-horizon restless multi-armed multi-action bandit problem and is provably hard to solve. To overcome this challenge, we propose a computationally appealing index policy entitled Quality Index Policy, which is well-defined without the Whittle indexability condition and is provably asymptotically optimal without the global attractor condition. These two conditions are widely needed in the design of most existing index policies, which are difficult to establish in general. Since the wireless access edge network environment is highly dynamic with system parameters unknown and time-varying, we further develop an index-aware reinforcement learning (RL) algorithm dubbed QA-UCB. We show that QA-UCB achieves a sub-linear regret with a low-complexity since it fully exploits the structure of the Quality Index Policy for making decisions. Extensive simulations using real-world traces demonstrate significant gains of proposed policies over conventional approaches. We note that the proposed framework for designing index policy and index-aware RL algorithm is of independent interest and could be useful for other large-scale multi-user problems.