Distributed TD(0) With Almost No Communication

Distributed TD(0) With Almost No Communication
复制标题

DOI:
10.1109/lcsys.2023.3287952
复制
发表时间:
2021-04
影响因子:
3
通讯作者:
R. Liu;Alexander Olshevsky
R. Liu;Alexander Olshevsky
中科院分区:
--
文献类型:
--
作者:
R. Liu;Alexander Olshevsky

文献摘要

相似文献

我们提供了一个新的非渐近分析的分布式时间差分学习与线性函数逼近。我们的方法依赖于“一次平均”,其中N个代理运行相同的TD(0)方法的本地副本,并在最后仅对结果进行一次平均。我们证明了一个版本的线性时间加速现象,其中的分布式过程的收敛时间是一个因素的N比TD(0)的收敛时间快。这是第一个结果证明从并行时间差分方法的好处。
We provide a new non-asymptotic analysis of distributed temporal difference learning with linear function approximation. Our approach relies on “one-shot averaging,” where N agents run identical local copies of the TD(0) method and average the outcomes only once at the very end. We demonstrate a version of the linear time speedup phenomenon, where the convergence time of the distributed process is a factor of N faster than the convergence time of TD(0). This is the first result proving benefits from parallelism for temporal difference methods.