Learning Good State and Action Representations via Tensor Decomposition
Learning Good State and Action Representations via Tensor Decomposition
复制标题
DOI:
10.1109/isit45174.2021.9518158
复制
发表时间:
2021-07
期刊:
影响因子:
--
通讯作者:
Chengzhuo Ni;Anru Zhang;Yaqi Duan;Mengdi Wang
中科院分区:
文献类型:
--
作者:
Chengzhuo Ni;Anru Zhang;Yaqi Duan;Mengdi Wang
The transition kernel of a continuous-state-action Markov decision process (MDP) admits a natural tensor structure. This paper proposes a tensor-inspired unsupervised learning method to identify meaningful low-dimensional state and action representations from empirical trajectories. The method exploits the MDP's tensor structure by kernelization, importance sampling and low-Tucker-rank approximation. This method can be further used to cluster states and actions respectively and find the best discrete MDP abstraction. We provide sharp statistical error bounds for tensor concentration and the preservation of diffusion distance after embedding.