Experience-driven Networking: A Deep Reinforcement Learning based Approach

Experience-driven Networking: A Deep Reinforcement Learning based Approach
复制标题

DOI:
10.1109/infocom.2018.8485853
复制
发表时间:
2018-01
期刊:
IEEE INFOCOM 2018 - IEEE Conference on Computer Communications
影响因子:
--
通讯作者:
Zhiyuan Xu;Jian Tang;Jingsong Meng;Weiyi Zhang;Yanzhi Wang;C. Liu;Dejun Yang
Zhiyuan Xu;Jian Tang;Jingsong Meng;Weiyi Zhang;Yanzhi Wang;C. Liu;Dejun Yang
中科院分区:
其他
文献类型:
--
作者:
Zhiyuan Xu;Jian Tang;Jingsong Meng;Weiyi Zhang;Yanzhi Wang;C. Liu;Dejun Yang

文献摘要

被引文献

相似文献

现代通信网络已经变得非常复杂和高度动态,这使得它们难以建模、预测和控制。在本文中,我们开发了一种新颖的经验驱动方法,可以从自己的经验而不是精确的数学模型中学习如何很好地控制通信网络,就像人类学习一项新技能(如驾驶、游泳等)一样。具体来说,我们首次提出利用新兴的深度强化学习(DRL)在通信网络中实现无模型控制;并提出了一种新颖且高效的基于drl的控制框架,DRL-TE,用于解决基本的网络问题:流量工程(TE)。该框架通过联合学习网络环境及其动态,并在强大的深度神经网络(Deep Neural Networks, dnn)的指导下进行决策,实现广泛使用的效用函数的最大化。我们提出了两种新技术,即TE感知探索和基于参与者关键的优先体验回放,以优化通用DRL框架,特别是针对TE。为了验证和评估所提出的框架,我们在ns-3中实现了它,并使用代表性和随机生成的网络拓扑对其进行了全面测试。大量的包级仿真结果表明:1)与几种广泛使用的基线方法相比,DRL-TE显著降低了端到端延迟,并不断提高网络效用,同时提供更好或相当的吞吐量;2) DRL-TE对网络变化具有鲁棒性;3) DRL- te始终优于最先进的DRL方法(用于连续控制),深度确定性策略梯度(DDPG),然而,后者不能提供令人满意的性能。
Modern communication networks have become very complicated and highly dynamic, which makes them hard to model, predict and control. In this paper, we develop a novel experience-driven approach that can learn to well control a communication network from its own experience rather than an accurate mathematical model, just as a human learns a new skill (such as driving, swimming, etc). Specifically, we, for the first time, propose to leverage emerging Deep Reinforcement Learning (DRL) for enabling model-free control in communication networks; and present a novel and highly effective DRL-based control framework, DRL-TE, for a fundamental networking problem: Traffic Engineering (TE). The proposed framework maximizes a widely-used utility function by jointly learning network environment and its dynamics, and making decisions under the guidance of powerful Deep Neural Networks (DNNs). We propose two new techniques, TE-aware exploration and actor-critic-based prioritized experience replay, to optimize the general DRL framework particularly for TE. To validate and evaluate the proposed framework, we implemented it in ns-3, and tested it comprehensively with both representative and randomly generated network topologies. Extensive packet-level simulation results show that 1) compared to several widely-used baseline methods, DRL-TE significantly reduces end-to-end delay and consistently improves the network utility, while offering better or comparable throughput; 2) DRL-TE is robust to network changes; and 3) DRL-TE consistently outperforms a state-of-the-art DRL method (for continuous control), Deep Deterministic Policy Gradient (DDPG), which, however, does not offer satisfying performance.