OnSlicing: online end-to-end network slicing with reinforcement learning

OnSlicing: online end-to-end network slicing with reinforcement learning
复制标题

DOI:
10.1145/3485983.3494850
复制
发表时间:
2021-11
期刊:
Proceedings of the 17th International Conference on emerging Networking EXperiments and Technologies
影响因子:
--
通讯作者:
Qiang Liu;Nakjung Choi;Tao Han
Qiang Liu;Nakjung Choi;Tao Han
中科院分区:
其他
文献类型:
--
作者:
Qiang Liu;Nakjung Choi;Tao Han

文献摘要

被引文献

相似文献

网络切片允许移动的网络运营商虚拟化基础设施并提供定制切片以支持具有异构需求的各种用例。在线深度强化学习(DRL)在解决网络问题和消除模拟与现实之间的差异方面显示出了很大的潜力。然而,使用在线DRL优化跨域资源是具有挑战性的,因为DRL的随机探索违反了切片的服务水平协议(SLA)和基础设施的资源约束。在本文中,我们提出了OnSlicing,一个在线的端到端的网络切片系统,以实现最小的资源使用,同时满足切片的SLA。OnSlicing允许对每个切片进行个性化学习,并通过使用一种新的约束感知策略更新方法和主动基线切换机制来维护其SLA。OnSlicing通过使用切片中的动作修改和基础设施中的参数协调的独特设计来遵守基础设施的资源约束。OnSlicing通过离线模仿基于规则的解决方案,进一步减轻了早期学习阶段在线学习的不良性能。此外,我们设计了四个新的域管理器,使动态资源配置在无线接入,传输,核心和边缘网络,分别在亚秒的时间尺度。我们在基于OpenAirInterface设计的端到端切片测试平台上实现了OnSlicing,该测试平台具有4G LTE和5G NR,OpenDayLight SDN平台和OpenAir-CN核心网络。实验结果表明,OnSlicing实现了61.3%的使用减少相比,基于规则的解决方案,并保持几乎为零违规(0.06%)在整个在线学习阶段。随着在线学习的融合,与最先进的在线DRL解决方案相比,OnSlicing减少了12.5%的使用量,而没有任何违规行为。
Network slicing allows mobile network operators to virtualize infrastructures and provide customized slices for supporting various use cases with heterogeneous requirements. Online deep reinforcement learning (DRL) has shown promising potential in solving network problems and eliminating the simulation-to-reality discrepancy. Optimizing cross-domain resources with online DRL is, however, challenging, as the random exploration of DRL violates the service level agreement (SLA) of slices and resource constraints of infrastructures. In this paper, we propose OnSlicing, an online end-to-end network slicing system, to achieve minimal resource usage while satisfying slices' SLA. OnSlicing allows individualized learning for each slice and maintains its SLA by using a novel constraint-aware policy update method and proactive baseline switching mechanism. OnSlicing complies with resource constraints of infrastructures by using a unique design of action modification in slices and parameter coordination in infrastructures. OnSlicing further mitigates the poor performance of online learning during the early learning stage by offline imitating a rule-based solution. Besides, we design four new domain managers to enable dynamic resource configuration in radio access, transport, core, and edge networks, respectively, at a timescale of subseconds. We implement OnSlicing on an end-to-end slicing testbed designed based on OpenAirInterface with both 4G LTE and 5G NR, OpenDayLight SDN platform, and OpenAir-CN core network. The experimental results show that OnSlicing achieves 61.3% usage reduction as compared to the rule-based solution and maintains nearly zero violation (0.06%) throughout the online learning phase. As online learning is converged, OnSlicing reduces 12.5% usage without any violations as compared to the state-of-the-art online DRL solution.