A Framework for Mapping DRL Algorithms With Prioritized Replay Buffer Onto Heterogeneous Platforms

A Framework for Mapping DRL Algorithms With Prioritized Replay Buffer Onto Heterogeneous Platforms
复制标题

DOI:
10.1109/tpds.2023.3264823
复制
发表时间:
2023-06
影响因子:
5.3
通讯作者:
Chi Zhang;Yuan Meng;V. Prasanna
Chi Zhang;Yuan Meng;V. Prasanna
中科院分区:
计算机科学2区
文献类型:
--
作者:
Chi Zhang;Yuan Meng;V. Prasanna

文献摘要

相似文献

尽管深度强化学习(DRL)最近在自动驾驶汽车、机器人和监控方面取得了成功,但训练DRL代理需要大量的时间和计算资源。在本文中,我们的目标是加速DRL与优先重放缓冲区,由于其在各种基准测试的最先进的性能。优先回放缓冲区DRL的计算原语包括环境仿真、神经网络推理、优先回放缓冲区采样、优先回放缓冲区更新和神经网络训练。运行这些原语的速度因各种DRL算法而异,例如深度Q网络和深度确定性策略梯度。这使得DRL算法的固定映射效率低下。在这项工作中,我们提出了一个框架DRL算法映射到异构平台组成的多核CPU,GPU和FPGA。首先,我们在CPU,FPGA和GPU上为每个原语开发特定的加速器。其次,我们放松优先级更新和采样之间的数据依赖性优先级重放缓冲区。通过这样做,GPU,FPGA和CPU之间的数据传输所造成的延迟可以完全隐藏,而不会牺牲使用目标DRL算法学习的代理所获得的回报。最后,给定DRL算法规范,我们的设计空间探索自动选择基于分析性能模型的各种原语的最佳映射。在广泛使用的基准测试环境中,我们的实验结果表明,与相同异构平台上的基线映射相比,训练吞吐量提高了997.3倍。与最先进的分布式强化学习框架RLlib相比,我们在训练吞吐量方面实现了1.06$\times \sim$× 1005×的提高。
Despite the recent success of Deep Reinforcement Learning (DRL) in self-driving cars, robotics and surveillance, training DRL agents takes tremendous amount of time and computation resources. In this article, we aim to accelerate DRL with Prioritized Replay Buffer due to its state-of-the-art performance on various benchmarks. The computation primitives of DRL with Prioritized Replay Buffer include environment emulation, neural network inference, sampling from Prioritized Replay Buffer, updating Prioritized Replay Buffer and neural network training. The speed of running these primitives varies for various DRL algorithms such as Deep Q Network and Deep Deterministic Policy Gradient. This makes a fixed mapping of DRL algorithms inefficient. In this work, we propose a framework for mapping DRL algorithms onto heterogeneous platforms consisting of a multi-core CPU, a GPU and a FPGA. First, we develop specific accelerators for each primitive on CPU, FPGA and GPU. Second, we relax the data dependency between priority update and sampling performed in the Prioritized Replay Buffer. By doing so, the latency caused by data transfer between GPU, FPGA and CPU can be completely hidden without sacrificing the rewards achieved by agents learned using the target DRL algorithms. Finally, given a DRL algorithm specification, our design space exploration automatically chooses the optimal mapping of various primitives based on an analytical performance model. On widely used benchmark environments, our experimental results demonstrate up to 997.3× improvement in training throughput compared with baseline mappings on the same heterogeneous platform. Compared with the state-of-the-art distributed Reinforcement Learning framework RLlib, we achieve 1.06$\times \sim$×∼ 1005× improvement in training throughput.