Genet: automatic curriculum generation for learning adaptation in networking

Genet: automatic curriculum generation for learning adaptation in networking
复制标题

DOI:
10.1145/3544216.3544243
复制
发表时间:
2022-02
期刊:
Proceedings of the ACM SIGCOMM 2022 Conference
影响因子:
--
通讯作者:
Zhengxu Xia;Yajie Zhou;Francis Y. Yan;Junchen Jiang
Zhengxu Xia;Yajie Zhou;Francis Y. Yan;Junchen Jiang
中科院分区:
其他
文献类型:
--
作者:
Zhengxu Xia;Yajie Zhou;Francis Y. Yan;Junchen Jiang

文献摘要

被引文献

相似文献

深度强化学习(RL)在展示其网络优势的同时,其缺陷也引起了公众的注意。在大范围的网络环境上进行训练会导致次优的性能,而在小范围的环境上进行训练会导致较差的泛化。这项工作提出了Genet,一个新的训练框架,用于学习更好的基于强化学习的网络自适应算法。Genet是建立在课程学习的基础上的,这在其他RL应用中已经被证明是有效的。在高水平上,课程学习逐渐为训练提供更多的“困难”环境,而不是随机选择统一的环境。然而,由于网络环境的“难度”是未知的,因此在网络环境中应用课程学习并非易事。我们的见解是利用传统的基于规则的(非RL)基线:如果当前的RL模型在网络环境中的表现明显不如基于规则的基线,那么在这种环境中进一步训练它往往会带来实质性的改进。Genet会自动搜索这样的环境,并迭代地将它们提升到训练中。三个案例研究——自适应视频流、拥塞控制和负载平衡——证明了Genet生成的RL策略优于常规训练的RL策略和传统基线。
As deep reinforcement learning (RL) showcases its strengths in networking, its pitfalls are also coming to the public's attention. Training on a wide range of network environments leads to suboptimal performance, whereas training on a narrow distribution of environments results in poor generalization. This work presents Genet, a new training framework for learning better RL-based network adaptation algorithms. Genet is built on curriculum learning, which has proved effective against similar issues in other RL applications. At a high level, curriculum learning gradually feeds more "difficult" environments to the training rather than choosing them uniformly at random. However, applying curriculum learning in networking is nontrivial since the "difficulty" of a network environment is unknown. Our insight is to leverage traditional rule-based (non-RL) baselines: If the current RL model performs significantly worse in a network environment than the rule-based baselines, then further training it in this environment tends to bring substantial improvement. Genet automatically searches for such environments and iteratively promotes them to training. Three case studies---adaptive video streaming, congestion control, and load balancing---demonstrate that Genet produces RL policies that outperform both regularly trained RL policies and traditional baselines.