Scalable Deep Reinforcement Learning-Based Online Routing for Multi-Type Service Requirements

Scalable Deep Reinforcement Learning-Based Online Routing for Multi-Type Service Requirements
复制标题

DOI:
10.1109/tpds.2023.3284651
复制
发表时间:
2023-08
影响因子:
5.3
通讯作者:
Chenyi Liu;Pingfei Wu;Mingwei Xu;Yuan Yang;Nan Geng
Chenyi Liu;Pingfei Wu;Mingwei Xu;Yuan Yang;Nan Geng
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chenyi Liu;Pingfei Wu;Mingwei Xu;Yuan Yang;Nan Geng

文献摘要

相似文献

新兴的应用程序对Internet提出了关键的QoS要求。流分类技术、软件定义网络(SDN)和可编程网络设备的改进使得快速识别用户需求和控制细粒度流量的路由成为可能。同时,对在线方式下具有多种QoS需求的流量流的转发路径优化问题没有得到充分的解决。为了解决这个问题,我们提出了DRL-OR-S,一种使用多智能体深度强化学习的高度可扩展的在线路由算法。DRL-OR-S采用综合奖励函数、高效的学习算法和新颖的深度神经网络结构,根据不同类型的流量需求学习合适的路由策略。为了增强泛化和可扩展性,我们提出了一种新的基于图的行动者-评论家网络架构和精心设计的DRL-OR-S输入状态。为了加速训练过程和保证可靠性,我们进一步引入了一个神经网络模拟器来进行有效的离线训练,并引入了一个安全的学习机制来避免在线路由过程中的不安全路由。我们在SDN架构下实现了DRL-OR-S,并利用真实的网络拓扑和流量轨迹进行了基于mininet的实验。结果表明,DRL-OR-S能很好地同时满足时延敏感、吞吐量敏感、时延-吞吐量敏感和时延-损失敏感流的需求,同时在链路故障、流量变化、不可见的大拓扑和部分部署场景下表现出很强的自适应能力和可靠性。
Emerging applications raise critical QoS requirements for the Internet. The improvements in flow classification technologies, software-defined networks (SDN), and programmable network devices make it possible to fast identify users’ requirements and control the routing for fine-grained traffic flows. Meanwhile, the problem of optimizing the forwarding paths for traffic flows with multiple QoS requirements in an online fashion is not addressed sufficiently. To address the problem, we propose DRL-OR-S, a highly scalable online routing algorithm using multi-agent deep reinforcement learning. DRL-OR-S adopts a comprehensive reward function, an efficient learning algorithm, and a novel deep neural network structure to learn appropriate routing strategies for different types of flow requirements. In order to enhance the generalization and scalability, we propose a novel graph-based actor-critic network architecture and a carefully designed input state for DRL-OR-S. To accelerate the training process and guarantee reliability, we further introduce an NN-simulator for efficient offline training and a safe learning mechanism to avoid unsafe routes during the online routing process. We implement DRL-OR-S under SDN architecture and conduct Mininet-based experiments using real network topologies and traffic traces. The results validate that DRL-OR-S can well satisfy the requirements of latency-sensitive, throughput-sensitive, latency-throughput-sensitive, and latency-loss-sensitive flows at the same time, while exhibiting great adaptiveness and reliability under the scenarios of link failure, traffic change, unseen large topology and partial deployment.