Distributed and Distribution-Robust Meta Reinforcement Learning (D$^{2}$-RMRL) for Data Pre-Storage and Routing in Cube Satellite Networks

Distributed and Distribution-Robust Meta Reinforcement Learning (D$^{2}$-RMRL) for Data Pre-Storage and Routing in Cube Satellite Networks
复制标题

DOI:
10.1109/jstsp.2022.3232944
复制
发表时间:
2022-06
影响因子:
7.5
通讯作者:
Ye Hu;Xiaodong Wang;W. Saad
Ye Hu;Xiaodong Wang;W. Saad
中科院分区:
工程技术1区
文献类型:
--
作者:
Ye Hu;Xiaodong Wang;W. Saad

文献摘要

相似文献

研究了动态、资源受限立方体卫星网络中的数据预存储和路由问题。在这样的网络中,每个立方体卫星将所请求的数据传送到其覆盖范围内的用户集群。一组地面网关将路由和预存储某些数据到卫星,以便地面用户可以直接使用预存储的数据。这个预存储和路由设计问题被公式化为分散式马尔可夫决策过程(Dec-MDP),在该过程中,我们寻求找到最大化预存储命中率的最优策略,即,用户的一部分直接被预存储的数据服务。为了获得最优策略,提出了一种分布式鲁棒Meta强化学习(D$^{2}$-RMRL)算法,该算法由三个关键部分组成:值分解,用于在分布式环境下以最小的通信开销获得全局最优; Meta学习,用于在动态条件下获得最优初始值,以减少训练时间;预训练,用于进一步加快Meta训练过程。仿真结果表明,使用所提出的值分解和Meta训练技术,卫星网络可以实现31.8%的提高预存储命中和40.7%的提高收敛速度,相比基线强化学习算法。此外,使用所提出的预训练机制有助于将元学习过程缩短高达43.7%。
In this paper, the problem of data pre-storage and routing in dynamic, resource-constrained cube satellite networks is studied. In such a network, each cube satellite delivers requested data to user clusters under its coverage. A group of ground gateways will route and pre-store certain data to the satellites, such that the ground users can be directly served with the pre-stored data. This pre-storage and routing design problem is formulated as a decentralized Markov decision process (Dec-MDP) in which we seek to find the optimal strategy that maximizes the pre-store hit rate, i.e., the fraction of users being directly served with the pre-stored data. To obtain the optimal strategy, a distributed distribution-robust meta reinforcement learning (D$^{2}$-RMRL) algorithm is proposed that consists of three key ingredients: value-decomposition for achieving the global optimum in distributed setting with minimum communication overhead, meta learning to obtain the optimal initial to reduce the training time under dynamic conditions, and pre-training to further speed up the meta training procedure. Simulation results show that, using the proposed value decomposition and meta training techniques, the satellite networks can achieve a 31.8% improvement of the pre-store hits and a 40.7% improvement of the convergence speed, compared to a baseline reinforcement learning algorithm. Moreover, the use of the proposed pre-training mechanism helps to shorten the meta-learning procedure by up to 43.7%.