Reinforcement Learning-Based Near-Optimal Load Balancing for Heterogeneous LiFi WiFi Network

Reinforcement Learning-Based Near-Optimal Load Balancing for Heterogeneous LiFi WiFi Network
复制标题

DOI:
10.1109/jsyst.2021.3088302
复制
发表时间:
2021-06
影响因子:
4.4
通讯作者:
Rizwana Ahmad;Mohammad Dehghani Soltani;M. Safari;A. Srivastava
Rizwana Ahmad;Mohammad Dehghani Soltani;M. Safari;A. Srivastava
中科院分区:
计算机科学2区
文献类型:
--
作者:
Rizwana Ahmad;Mohammad Dehghani Soltani;M. Safari;A. Srivastava

文献摘要

被引文献

相似文献

由于光谱不重叠,光保真(LiFi)和WiFi技术可以共存并形成异构LiFi WiFi网络(HLWN)。HLWN的性能很大程度上取决于负载均衡策略。由于HLWN的负载均衡是一个非凸的混合整数非线性规划优化问题,在数学上是难以处理的,因此,传统的优化方法无法提供全局最优解。虽然可以使用穷举搜索方法来获得最优解,但是这在计算上是复杂的。因此,在这篇文章中,强化学习(RL)为基础的算法进行了探索,以合理的低复杂度和接近最佳性能的下行链路HLWN的负载平衡问题。我们为RL提出了三种不同的奖励函数;第一种和第二种奖励函数分别用于最大化平均网络吞吐量和用户满意度。第三个奖励函数是为了最大化长期系统吞吐量,并确保所有用户至少50%的用户满意度。为了研究链路聚合对系统性能的影响,本文考虑了两种不同类型的接收机方案,即单接入点(SAP)和链路聚合(LA)方案。虽然SAP允许用户仅从SAP接收数据,但LA方案允许用户同时从LiFi和WiFi AP接收数据。本文还包括接收机设备的随机方向和切换开销的影响。此外,领域知识的概念已经在这篇文章中,以减少算法的计算复杂度。该系统的性能进行了比较,以下两个基准:接收信号强度(RSS)和穷举搜索的计算复杂度,平均系统吞吐量和用户满意度的基础上。结果表明,建议的RL计划优于RSS计划的平均系统吞吐量和用户满意度。具有适当奖励函数的RL方案以合理低的复杂度提供与穷举搜索匹配的性能。
Owing to the nonoverlapping spectrum, light fidelity (LiFi) and WiFi technologies can coexist and form a heterogeneous LiFi WiFi network (HLWN). The performance of HLWN significantly depends upon the load balancing strategies. Since load balancing of HLWN is a nonconvex mixed-integer nonlinear programming optimization problem, it is mathematically intractable, and therefore, the conventional optimization methods fail to provide an optimal global solution. Although an optimal solution can be obtained using the exhaustive search method, it would be computationally complex. Therefore, in this article, a reinforcement learning (RL)-based algorithm is explored for solving the load balancing problem for the downlink HLWN at reasonably low complexity and near optimal performance. We have proposed three different reward functions for RL; the first and second reward functions work toward maximizing average network throughput and user satisfaction, respectively. The third reward function is designed to maximize the long-term system throughput and ensure at least 50% user’s satisfaction for all users. In order to study the effects of link aggregation on the system performance, this article considers two different types of receiver schemes, namely, single access point (SAP) and link aggregation (LA) scheme. While the SAP allows the user to receive data only from an SAP, the LA scheme allows the user to receive data simultaneously from both LiFi and WiFi AP. This article also includes effect of random orientation of the receiver device and handover overhead. Furthermore, concepts of domain knowledge have been included in this article to reduce the computational complexity of the algorithm. The proposed system performance is compared with the the following two benchmarks: received signal strength (RSS) and exhaustive search based on the computational complexity, average system throughput, and user satisfaction. It is shown that the proposed RL scheme outperforms the RSS scheme in average system throughput and user satisfaction. The RL scheme with an appropriate reward function provides a matching performance to the exhaustive search at reasonably low complexity.