DeepRMSA: A Deep Reinforcement Learning Framework for Routing, Modulation and Spectrum Assignment in Elastic Optical Networks

DeepRMSA: A Deep Reinforcement Learning Framework for Routing, Modulation and Spectrum Assignment in Elastic Optical Networks
复制标题

DOI:
10.1109/jlt.2019.2923615
复制
发表时间:
2019-08-15
影响因子:
4.7
通讯作者:
Yoo, S. J. Ben
Yoo, S. J. Ben
中科院分区:
工程技术2区
文献类型:
--
作者:
Chen, Xiaoliang;Li, Baojia;Yoo, S. J. Ben

文献摘要

被引文献

相似文献

本文提出了DeepRMSA,这是一种用于弹性光网络(EONs)中路由、调制和频谱分配(RMSA)的深度强化学习框架。DeepRMSA通过使用可以感知复杂EON状态的深度神经网络(DNN)参数化策略来学习正确的在线RMSA策略。DNN通过动态光路供应的经验进行训练。我们首先修改了异步优势演员-评论家算法,并提出了一种基于情节的DeepRMSA训练机制,即DeepRMSA-EP。DeepRMSA-EP将动态配置过程划分为多个片段(每个片段包含固定数量的光路请求的服务),并在每个片段结束时执行训练。DeepRMSA-EP在服务请求的每一步的优化目标是最大化剩余片段中的累积奖励。因此,我们需要估计未知的未来状态相关的奖励。为了克服DeepRMSA-EP训练中由于累积奖励的振荡而导致的不稳定问题,我们进一步提出了一种基于窗口的灵活训练机制,即,DeepRMSA-FLX. DeepRMSA-FLX试图通过将每一步的优化范围定义为滑动窗口来消除振荡,并确保累积奖励始终包括来自固定数量请求的奖励。两个样本拓扑的评估表明,与基线相比,DeepRMSA-FLX可以有效地稳定训练,同时实现阻塞概率降低超过20.3%和14.3%。
This paper proposes DeepRMSA, a deep reinforcement learning framework for routing, modulation and spectrum assignment (RMSA) in elastic optical networks (EONs). DeepRMSA learns the correct online RMSA policies by parameterizing the policies with deep neural networks (DNNs) that can sense complex EON states. The DNNs are trained with experiences of dynamic lightpath provisioning. We first modify the asynchronous advantage actor-critic algorithm and present an episode-based training mechanism for DeepRMSA, namely, DeepRMSA-EP. DeepRMSA-EP divides the dynamic provisioning process into multiple episodes (each containing the servicing of a fixed number of lightpath requests) and performs training by the end of each episode. The optimization target of DeepRMSA-EP at each step of servicing a request is to maximize the cumulative reward within the rest of the episode. Thus, we obviate the need for estimating the rewards related to unknown future states. To overcome the instability issue in the training of DeepRMSA-EP due to the oscillations of cumulative rewards, we further propose a window-based flexible training mechanism, i.e., DeepRMSA-FLX. DeepRMSA-FLX attempts to smooth out the oscillations by defining the optimization scope at each step as a sliding window, and ensuring that the cumulative rewards always include rewards from a fixed number of requests. Evaluations with the two sample topologies show that DeepRMSA-FLX can effectively stabilize the training while achieving blocking probability reductions of more than 20.3% and 14.3%, when compared with the baselines.