Distributed Reinforcement Learning for Age of Information Minimization in Real-Time IoT Systems

Distributed Reinforcement Learning for Age of Information Minimization in Real-Time IoT Systems
复制标题

DOI:
10.1109/jstsp.2022.3144874
复制
发表时间:
2021-04
影响因子:
7.5
通讯作者:
Sihua Wang;Mingzhe Chen;Zhaohui Yang;Changchuan Yin;W. Saad;Shuguang Cui;H. Poor
Sihua Wang;Mingzhe Chen;Zhaohui Yang;Changchuan Yin;W. Saad;Shuguang Cui;H. Poor
中科院分区:
工程技术1区
文献类型:
--
作者:
Sihua Wang;Mingzhe Chen;Zhaohui Yang;Changchuan Yin;W. Saad;Shuguang Cui;H. Poor

文献摘要

被引文献

相似文献

本文研究了最小化物联网(IoT)设备信息年龄(AoI)和总能耗的加权和的问题。在所考虑的模型中,每个物联网设备都会监控遵循非线性动力学的物理过程。由于物理过程的动态随时间变化,每个设备应该找到最佳采样频率来对物理系统的实时动态进行采样,并将采样信息发送到基站(BS)。由于无线资源有限,BS只能选择设备的子集来传输它们的采样信息。因此,边缘设备可以根据本地观测结果协同采样其监测到的动态,BS将立即从设备收集采样信息,从而避免用于采样和信息传输的额外时间和能量。为此,需要共同优化各设备的采样策略和基站的设备选择方案,以最小的能量准确监测物理过程的动态。该问题被表述为一个优化问题,其目标是最小化 AoI 成本和能耗的加权和。为了解决这个问题,我们提出了一种新颖的分布式强化学习(RL)方法来进行采样策略优化。所提出的算法使边缘设备能够使用自己的本地观察来协作找到全局最优采样策略。给定采样策略,可以优化设备选择方案,从而最小化所有设备的 AoI 和能耗的加权和。对真实PM 2.5污染数据的模拟表明,与传统的深度Q网络方法和均匀采样策略相比,该算法可以将AoI之和分别降低高达17.8%和33.9%,总能耗分别降低高达13.2%和35.1%。
In this paper, the problem of minimizing the weighted sum of age of information (AoI) and total energy consumption of Internet of Things (IoT) devices is studied. In the considered model, each IoT device monitors a physical process that follows nonlinear dynamics. As the dynamics of the physical process vary over time, each device should find an optimal sampling frequency to sample the real-time dynamics of the physical system and send sampled information to a base station (BS). Due to limited wireless resources, the BS can only select a subset of devices to transmit their sampled information. Thus, edge devices can cooperatively sample their monitored dynamics based on the local observations and the BS will collect the sampled information from the devices immediately, hence avoiding the additional time and energy used for sampling and information transmission. To this end, it is necessary to jointly optimize the sampling policy of each device and the device selection scheme of the BS so as to accurately monitor the dynamics of the physical process using minimum energy. This problem is formulated as an optimization problem whose goal is to minimize the weighted sum of AoI cost and energy consumption. To solve this problem, we propose a novel distributed reinforcement learning (RL) approach for the sampling policy optimization. The proposed algorithm enables edge devices to cooperatively find the global optimal sampling policy using their own local observations. Given the sampling policy, the device selection scheme can be optimized thus minimizing the weighted sum of AoI and energy consumption of all devices. Simulations with real PM 2.5 pollution data show that the proposed algorithm can reduce the sum of AoI by up to 17.8% and 33.9%, respectively, and the total energy consumption by up to 13.2% and 35.1%, respectively, compared to a conventional deep Q network method and a uniform sampling policy.