Distributed Online System Identification for LTI Systems Using Reverse Experience Replay

Distributed Online System Identification for LTI Systems Using Reverse Experience Replay
复制标题

DOI:
10.1109/cdc51059.2022.9992456
复制
发表时间:
2022-07
期刊:
2022 IEEE 61st Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Ting-Jui Chang;Shahin Shahrampour
Ting-Jui Chang;Shahin Shahrampour
中科院分区:
其他
文献类型:
--
作者:
Ting-Jui Chang;Shahin Shahrampour

文献摘要

相似文献

线性时不变(LTI)系统的辨识在控制和强化学习中起着重要作用。渐近和有限时间离线系统辨识在文献中得到了很好的研究。对于在线系统辨识,最近提出了具有反向经验重放的随机梯度下降(SGD-RER)的思想,其中数据序列存储在多个缓冲器中,并且随机梯度下降(SGD)更新在每个缓冲器中向后执行,以打破数据点之间的时间依赖性。受这项工作的启发,我们研究了分布式在线系统识别的LTI系统在多智能体网络。我们认为代理商是相同的LTI系统,网络的目标是通过利用代理商之间的通信来联合估计系统参数。我们提出了DSGD-RER,一个分布式的SGD-RER算法的变种,并从理论上描述了改进的估计误差相对于网络的大小。我们的数值实验证明,随着网络规模的增长,估计误差的减少。
Identification of linear time-invariant (LTI) systems plays an important role in control and reinforcement learning. Both asymptotic and finite-time offline system identification are well-studied in the literature. For online system identification, the idea of stochastic-gradient descent with reverse experience replay (SGD-RER) was recently proposed, where the data sequence is stored in several buffers and the stochastic-gradient descent (SGD) update performs backward in each buffer to break the time dependency between data points. Inspired by this work, we study distributed online system identification of LTI systems over a multi-agent network. We consider agents as identical LTI systems, and the network goal is to jointly estimate the system parameters by leveraging the communication between agents. We propose DSGD-RER, a distributed variant of the SGD-RER algorithm, and theoretically characterize the improvement of the estimation error with respect to the network size. Our numerical experiments certify the reduction of estimation error as the network size grows.