Learning an adaptive forwarding strategy for mobile wireless networks: resource usage vs. latency

Learning an adaptive forwarding strategy for mobile wireless networks: resource usage vs. latency
复制标题

DOI:
10.1007/s10994-024-06601-3
复制
发表时间:
2024-08-07
期刊:
影响因子:
7.5
通讯作者:
Wang,Bing
Wang,Bing
中科院分区:
计算机科学3区
文献类型:
--
作者:
Manfredi,Victoria;Wolfe,Alicia P.;Wang,Bing

文献摘要

相似文献

移动的无线网络对任何学习系统提出了若干挑战,这是由于设备移动的不确定性和可变性、分散的网络架构以及对网络资源的约束。在这项工作中,我们使用深度强化学习(DRL)来学习此类网络的可扩展和可推广的转发策略。我们做出了以下贡献:(i)我们使用分层RL来设计DRL分组代理而不是设备代理,以捕获随着时间的推移而做出的分组转发决策并提高训练效率;(ii)我们使用关系特征来确保学习的转发策略对广泛的网络动态的泛化性并实现离线训练;以及(iii)通过设计加权奖励函数,将转发目标和网络资源考虑结合到分组决策中。我们的研究结果表明,我们的DRL数据包代理使用的转发策略往往实现了类似的延迟每包交付的Oracle转发策略,几乎总是优于所有其他策略(包括国家的最先进的战略)在延迟方面,即使在DRL代理没有受过训练的情况下。
Mobile wireless networks present several challenges for any learning system, due to uncertain and variable device movement, a decentralized network architecture, and constraints on network resources. In this work, we use deep reinforcement learning (DRL) to learn a scalable and generalizable forwarding strategy for such networks. We make the following contributions: (i) we use hierarchical RL to design DRL packet agents rather than device agents to capture the packet forwarding decisions that are made over time and improve training efficiency; (ii) we use relational features to ensure generalizability of the learned forwarding strategy to a wide range of network dynamics and enable offline training; and (iii) we incorporate both forwarding goals and network resource considerations into packet decision-making by designing a weighted reward function. Our results show that the forwarding strategy used by our DRL packet agent often achieves a similar delay per packet delivered as the oracle forwarding strategy and almost always outperforms all other strategies (including state-of-the-art strategies) in terms of delay, even on scenarios on which the DRL agent was not trained.