Learning Good State and Action Representations for Markov Decision Process via Tensor Decomposition

Learning Good State and Action Representations for Markov Decision Process via Tensor Decomposition
复制标题

DOI:
--
复制
发表时间:
2021-05
期刊:
J. Mach. Learn. Res.
影响因子:
--
通讯作者:
Chengzhuo Ni;Yaqi Duan;M. Dahleh;Mengdi Wang;Anru R. Zhang
Chengzhuo Ni;Yaqi Duan;M. Dahleh;Mengdi Wang;Anru R. Zhang
中科院分区:
其他
文献类型:
--
作者:
Chengzhuo Ni;Yaqi Duan;M. Dahleh;Mengdi Wang;Anru R. Zhang

文献摘要

相似文献

连续状态-动作马尔可夫决策过程(MDP)的转移核具有自然张量结构。本文提出了一种张量启发的无监督学习方法,从经验轨迹中识别有意义的低维状态和动作表征。该方法通过核化、重要抽样和低Tucker秩近似地利用了MDP的张量结构。该方法还可以用来分别对状态和动作进行聚类,找到最优的离散MDP抽象。我们给出了张量集中和嵌入后扩散距离保持的尖锐的统计误差界。我们进一步证明了所学习的状态/动作抽象提供了对潜在块结构的精确近似(如果它们存在的话),从而使得在诸如策略评估之类的下游任务中能够进行函数近似。
The transition kernel of a continuous-state-action Markov decision process (MDP) admits a natural tensor structure. This paper proposes a tensor-inspired unsupervised learning method to identify meaningful low-dimensional state and action representations from empirical trajectories. The method exploits the MDP's tensor structure by kernelization, importance sampling and low-Tucker-rank approximation. This method can be further used to cluster states and actions respectively and find the best discrete MDP abstraction. We provide sharp statistical error bounds for tensor concentration and the preservation of diffusion distance after embedding. We further prove that the learned state/action abstractions provide accurate approximations to latent block structures if they exist, enabling function approximation in downstream tasks such as policy evaluation.