Learning Good State and Action Representations via Tensor Decomposition

Learning Good State and Action Representations via Tensor Decomposition
复制标题

DOI:
10.1109/isit45174.2021.9518158
复制
发表时间:
2021-07
期刊:
2021 IEEE International Symposium on Information Theory (ISIT)
影响因子:
--
通讯作者:
Chengzhuo Ni;Anru Zhang;Yaqi Duan;Mengdi Wang
Chengzhuo Ni;Anru Zhang;Yaqi Duan;Mengdi Wang
中科院分区:
其他
文献类型:
--
作者:
Chengzhuo Ni;Anru Zhang;Yaqi Duan;Mengdi Wang

文献摘要

相似文献

连续状态动作马尔可夫决策过程(MDP)的转移核承认自然张量结构。本文提出了一种受张量启发的无监督学习方法,用于从经验轨迹中识别有意义的低维状态和动作表示。该方法通过核化、重要性采样和低塔克秩近似来利用 MDP 的张量结构。该方法可以进一步用于分别对状态和动作进行聚类,并找到最佳的离散 MDP 抽象。我们为张量浓度和嵌入后扩散距离的保留提供了清晰的统计误差范围。
The transition kernel of a continuous-state-action Markov decision process (MDP) admits a natural tensor structure. This paper proposes a tensor-inspired unsupervised learning method to identify meaningful low-dimensional state and action representations from empirical trajectories. The method exploits the MDP's tensor structure by kernelization, importance sampling and low-Tucker-rank approximation. This method can be further used to cluster states and actions respectively and find the best discrete MDP abstraction. We provide sharp statistical error bounds for tensor concentration and the preservation of diffusion distance after embedding.