Deep Reinforcement Learning for Scheduling in Multi-Hop Wireless Networks : Invited Paper

Deep Reinforcement Learning for Scheduling in Multi-Hop Wireless Networks : Invited Paper
复制标题

DOI:
10.1109/mass52906.2021.00010
复制
发表时间:
2021-10
期刊:
2021 IEEE 18th International Conference on Mobile Ad Hoc and Smart Systems (MASS)
影响因子:
--
通讯作者:
Shuai Zhang;Bo Yin;Yu Cheng
Shuai Zhang;Bo Yin;Yu Cheng
中科院分区:
其他
文献类型:
--
作者:
Shuai Zhang;Bo Yin;Yu Cheng

文献摘要

相似文献

具有一定优化目标的无线网络中传输链接的有效调度,并受到干扰和网络流程约束在无线网络研究中起着核心作用。作为传统数学分析的替代方法,数据驱动的学习方法通​​过从无线网络中的机器学习的经验和灵感应用中提取知识来解决难题,从而解决了有望。在本文中,我们专注于通过机器学习来解决多跳无线网络中的基本调度问题,面临着非差异性操作的巨大挑战,并考虑了可变网络拓扑。为了解决这些问题,我们提出了一种基于加强学习的方法,以解决协议干扰模型下的一类网络流问题。从经验中学习,所提出的方法制定了一种策略,以顺序选择链接的最佳子集以同时传输以最大化系统吞吐量而不会引起干扰。该模型结构的设计方式包含网络拓扑信息,以允许灵活数量的网络节点,并允许非不同的决策操作传递信息性梯度信息。合成和现实部署数据的实验表明,所提出的算法以显着降低的时间成本接近最低的性能。
The efficient scheduling of transmission links in a wireless network with a certain optimization objective and subject to the interference and network flow constraints plays a central role in wireless networking research. As an alternative to traditional mathematical analysis, data-driven learning methods have shown promise in solving difficult problems by extracting knowledge from experiences and inspired applications of machine learning in wireless networking. In this paper, we focus on tackling the fundamental scheduling issue in multi-hop wireless networks with machine learning, facing the great challenges of the involvement of non-differentiable operations and the consideration of variable network topologies. To address these issues, we propose a reinforcement learning-based method to solve a class of network flow problems under the protocol interference model. Learning from experience, the proposed approach develops a strategy to sequentially select optimum subsets of links to transmit simultaneously to maximize the system throughput without causing interference. The model structure is designed in a way that incorporates network topological information to allow a flexible number of network nodes, and allows non-differentiable decision operation to pass informative gradient information. Experiments with synthetic and real-world deployment data demonstrate that the proposed algorithm achieves close-to-optimum performance at a significantly reduced time cost.