Tensor optimization with group lasso for multi-agent predictive state representation

Tensor optimization with group lasso for multi-agent predictive state representation
复制标题

DOI:
10.1016/j.knosys.2021.106893
复制
发表时间:
2021-03
期刊:
Knowl. Based Syst.
影响因子:
--
通讯作者:
Biyang Ma;Jing Tang;Bilian Chen;Yinghui Pan;Yi-feng Zeng
Biyang Ma;Jing Tang;Bilian Chen;Yinghui Pan;Yi-feng Zeng
中科院分区:
其他
文献类型:
--
作者:
Biyang Ma;Jing Tang;Bilian Chen;Yinghui Pan;Yi-feng Zeng

文献摘要

相似文献

预测状态表示(PSR)是动态系统的一种紧凑模型,它将状态表示为对未来可观测事件的预测向量。它是一种替代部分可观测马尔可夫决策过程(POMDP)模型在处理不确定性下的顺序决策问题。现有的PSR研究大多集中在单智能体环境下的模型学习。在本文中,我们研究了一个多智能体PSR模型上可用的智能体交互数据。事实证明,学习多智能体PSR模型是相当困难的,特别是在有限的样本和越来越多的智能体的情况下。我们诉诸张量技术,以更好地表示动态系统的特性,并解决具有挑战性的任务,学习多智能体PSR问题的基础上张量优化。我们首先专注于两个代理的情况下,并使用三阶张量(系统动力学张量)来捕获系统的交互数据。然后,PSR模型的发现可以被表述为一个张量优化问题与组套索,和交替方向的乘子方法被称为解决嵌入的子问题。因此,预测参数和状态向量可以直接从优化解中学习,并且过渡参数可以经由线性回归导出。随后,我们将张量学习方法推广到一个多(N> 2)主体PSR模型中,并分析了学习算法的计算复杂度。实验结果表明,张量优化方法在多问题域上学习多智能体PSR模型时具有良好的性能。
Predictive state representation (PSR) is a compact model of dynamic systems that represents state as a vector of predictions about future observable events. It is an alternative to a partially observable Markov decision process (POMDP) model in dealing with a sequential decision-making problem under uncertainty. Most of the existing PSR research focuses on the model learning in a single-agent setting. In this paper, we investigate a multi-agent PSR model upon available agents interaction data. It turns out to be rather difficult to learn a multi-agent PSR model especially with limited samples and increasing number of agents. We resort to a tensor technique to better represent dynamic system characteristics and address the challenging task of learning multi-agent PSR problems based on tensor optimization. We first focus on a two-agent scenario and use a third order tensor (system dynamics tensor) to capture the system interaction data. Then, the PSR model discovery can be formulated as a tensor optimization problem with group lasso, and an alternating direction method of multipliers is called for solving the embedded subproblems. Hence, the prediction parameters and state vectors can be directly learned from the optimization solutions, and the transition parameters can be derived via a linear regression. Subsequently, we generalize the tensor learning approach in a multi (N> 2)-agent PSR model, and analyze the computational complexity of the learning algorithms. Experimental results show that the tensor optimization approaches have provided promising performances on learning a multi-agent PSR model over multiple problem domains.