Experience-Driven Congestion Control: When Multi-Path TCP Meets Deep Reinforcement Learning

Experience-Driven Congestion Control: When Multi-Path TCP Meets Deep Reinforcement Learning
复制标题

DOI:
10.1109/jsac.2019.2904358
复制
发表时间:
2019-03
影响因子:
16.4
通讯作者:
Zhiyuan Xu;Jian Tang;Chengxiang Yin;Yanzhi Wang;G. Xue
Zhiyuan Xu;Jian Tang;Chengxiang Yin;Yanzhi Wang;G. Xue
中科院分区:
计算机科学1区
文献类型:
--
作者:
Zhiyuan Xu;Jian Tang;Chengxiang Yin;Yanzhi Wang;G. Xue

文献摘要

被引文献

相似文献

在本文中,我们的目标是通过利用新兴的深度学习从一个全新的角度来研究网络问题,开发一种经验驱动的方法,使网络或协议能够根据自己的经验(例如,运行时统计数据)学习控制自身的最佳方式,就像人类学习技能一样。给出了一个基于深度强化学习(DRL)的拥塞控制框架DRL-CC的设计、实现和评估,该框架实现了我们的经验驱动的多路径TCP拥塞控制设计思想。DRL-CC利用单个(而不是多个独立的)代理来动态地、联合地对终端主机上的所有活动MPTCP流执行拥塞控制,以最大化整体效用为目标。我们设计的新颖之处在于在DRL框架下利用灵活的递归神经网络LSTM来学习所有活动流的表示并处理它们的动态。此外,我们首次将上述基于LSTM的表示网络集成到用于连续(拥塞)控制的参与者-批评者框架中,该框架利用新出现的确定性策略梯度以端到端的方式训练批评者、参与者和LSTM网络。我们在Linux内核的MPTCP实现的基础上实现了DRL-CC。实验结果表明:1)在不牺牲公平性的前提下,DRL-CC在吞吐量方面始终显著优于几种著名的MPTCP拥塞控制算法;2)对于具有时变流量的高动态网络环境,DRL-CC具有良好的灵活性和健壮性;3)它与常规的TCP协议相比是友好的。
In this paper, we aim to study networking problems from a whole new perspective by leveraging emerging deep learning, to develop an experience-driven approach, which enables a network or a protocol to learn the best way to control itself from its own experience (e.g., runtime statistics data), just as a human learns a skill. We present design, implementation and evaluation of a deep reinforcement learning (DRL)-based control framework, DRL-CC (DRL for Congestion Control), which realizes our experience-driven design philosophy on multi-path TCP (MPTCP) congestion control. DRL-CC utilizes a single (instead of multiple independent) agent to dynamically and jointly perform congestion control for all active MPTCP flows on an end host with the objective of maximizing the overall utility. The novelty of our design is to utilize a flexible recurrent neural network, LSTM, under a DRL framework for learning a representation for all active flows and dealing with their dynamics. Moreover, we, for the first time, integrate the above LSTM-based representation network into an actor-critic framework for continuous (congestion) control, which leverages the emerging deterministic policy gradient to train critic, actor, and LSTM networks in an end-to-end manner. We implemented DRL-CC based on the MPTCP implementation in the Linux kernel. The experimental results show that 1) DRL-CC consistently and significantly outperforms a few well-known MPTCP congestion control algorithms in terms of goodput without sacrificing fairness, 2) it is flexible and robust to highly-dynamic network environments with time-varying flows, and 3) it is friendly to regular TCP.