Reinforcement Learning for Adaptive Resource Allocation in Fog RAN for IoT With Heterogeneous Latency Requirements

Reinforcement Learning for Adaptive Resource Allocation in Fog RAN for IoT With Heterogeneous Latency Requirements
复制标题

DOI:
10.1109/access.2019.2939735
复制
发表时间:
2019-09
期刊:
影响因子:
3.9
通讯作者:
A. Nassar;Yasin Yılmaz
A. Nassar;Yasin Yılmaz
中科院分区:
计算机科学3区
文献类型:
--
作者:
A. Nassar;Yasin Yılmaz

文献摘要

被引文献

相似文献

鉴于物联网(IoT)设备和应用的快速增长,最近已经提出了用于第五代(5G)无线通信的雾无线电接入网络(Fog-RAN),以确保无法适应大延迟的IoT应用的超可靠低延迟通信(URLLC)的要求。为此,雾节点(FN)配备了计算、信号处理和存储功能,将云的固有操作和服务扩展到边缘。我们考虑顺序分配FN的有限资源的异构延迟要求的物联网应用程序的问题。对于来自IoT用户的每个访问请求,FN需要决定是利用其自己的资源在边缘处本地服务它,还是将其提交给云以保存其有价值的资源用于对系统具有潜在更高效用的未来用户(即,较低的等待时间要求)。我们制定了一个马尔可夫决策过程(MDP)的形式的Fog-RAN资源分配问题,并采用几种强化学习(RL)方法,即Q学习,SARSA,预期SARSA和蒙特卡洛,通过学习最佳决策策略来解决MDP问题。我们验证了RL方法的性能和自适应性,并将其与具有各种切片阈值的网络切片方法的性能进行比较。考虑到19个具有异构延迟要求的物联网环境的广泛模拟结果证实,无论物联网环境如何,RL方法总是能够实现最佳性能。
In light of the quick proliferation of Internet of things (IoT) devices and applications, fog radio access network (Fog-RAN) has been recently proposed for fifth generation (5G) wireless communications to assure the requirements of ultra-reliable low-latency communication (URLLC) for the IoT applications which cannot accommodate large delays. To this end, fog nodes (FNs) are equipped with computing, signal processing and storage capabilities to extend the inherent operations and services of the cloud to the edge. We consider the problem of sequentially allocating the FN’s limited resources to IoT applications of heterogeneous latency requirements. For each access request from an IoT user, the FN needs to decide whether to serve it locally at the edge utilizing its own resources or to refer it to the cloud to conserve its valuable resources for future users of potentially higher utility to the system (i.e., lower latency requirement). We formulate the Fog-RAN resource allocation problem in the form of a Markov decision process (MDP), and employ several reinforcement learning (RL) methods, namely Q-learning, SARSA, Expected SARSA, and Monte Carlo, for solving the MDP problem by learning the optimum decision-making policies. We verify the performance and adaptivity of the RL methods and compare it with the performance of the network slicing approach with various slicing thresholds. Extensive simulation results considering 19 IoT environments of heterogeneous latency requirements corroborate that RL methods always achieve the best possible performance regardless of the IoT environment.