On the Power of Multitask Representation Learning in Linear MDP

On the Power of Multitask Representation Learning in Linear MDP
复制标题

论线性 MDP 中多任务表示学习的威力

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
S. Du
S. Du
中科院分区:
--
文献类型:
--
作者:
Rui Lu;Gao Huang;S. Du

文献摘要

参考文献

被引文献

相似文献

尽管多任务表示学习已成为强化学习(RL)的一种流行方法,但对为什么和何时起作用的理论理解仍然有限。本文介绍了在生成模型下,在线性马尔可夫决策过程(MDP)中多任务表示学习的统计益处的分析。在本文中,我们考虑了一个代理来学习表示函数$ \ phi $从函数类$ \ phi $ from $ t $ phi $从$ t $ source任务中,每项任务$ n $ data,然后使用学习的$ \ hat {\ phi} $以减少新任务所需的样品数量。我们首先发现一个\ emph {最不活化的feature-abundance}(lafa)标准,称为$ \ kappa $,我们证明了一种直接的最小二乘算法可以学习$ \ tilde {o}(o}(o}(o}(o}(o}(o}), h^2 \ sqrt {\ frac {\ mathcal {c}(\ phi)^2 \ kappa d} {nt}+\ frac {\ kappa d} {n}})$ sub-optimal。这里$ h $是计划范围,$ \ mathcal {c}(\ phi)$是$ \ phi $的复杂度度量,$ d $是表示的维度(通常$ d \ ll \ mathcal {c} (\ phi)$)和$ n $是新任务的样本数量。因此,所需的$ n $是$ o(\ kappa d h^4),用于接近零的次级优先级,它比$ o小得多(\ mathcal {c}(\ phi)(\ phi)^2 \ kappa d h^4)在设置中没有多任务表示学习,其亚典型性差距是$ \ tilde {o}(h^2 \ sqrt {\ frac {\ kappa \ mathcal {c} {c}(\ phi)^2d} {n}}})$。从理论上讲,这说明了多任务表示学习在降低样本复杂性方面的力量。此外,我们注意到,为了确保较高的样本效率,LAFA标准$ \ kappa $应该很小。实际上,$ \ kappa $的数量范围很大,具体取决于新任务的不同采样分布。这表明自适应抽样技术对于使$ \ kappa $仅取决于$ d $很重要。最后,我们提供了嘈杂的网格世界环境的经验结果,以证实我们的理论发现。
While multitask representation learning has become a popular approach in reinforcement learning (RL), theoretical understanding of why and when it works remains limited. This paper presents analyses for the statistical benefit of multitask representation learning in linear Markov Decision Process (MDP) under a generative model. In this paper, we consider an agent to learn a representation function $\phi$ out of a function class $\Phi$ from $T$ source tasks with $N$ data per task, and then use the learned $\hat{\phi}$ to reduce the required number of sample for a new task. We first discover a \emph{Least-Activated-Feature-Abundance} (LAFA) criterion, denoted as $\kappa$, with which we prove that a straightforward least-square algorithm learns a policy which is $\tilde{O}(H^2\sqrt{\frac{\mathcal{C}(\Phi)^2 \kappa d}{NT}+\frac{\kappa d}{n}})$ sub-optimal. Here $H$ is the planning horizon, $\mathcal{C}(\Phi)$ is $\Phi$'s complexity measure, $d$ is the dimension of the representation (usually $d\ll \mathcal{C}(\Phi)$) and $n$ is the number of samples for the new task. Thus the required $n$ is $O(\kappa d H^4)$ for the sub-optimality to be close to zero, which is much smaller than $O(\mathcal{C}(\Phi)^2\kappa d H^4)$ in the setting without multitask representation learning, whose sub-optimality gap is $\tilde{O}(H^2\sqrt{\frac{\kappa \mathcal{C}(\Phi)^2d}{n}})$. This theoretically explains the power of multitask representation learning in reducing sample complexity. Further, we note that to ensure high sample efficiency, the LAFA criterion $\kappa$ should be small. In fact, $\kappa$ varies widely in magnitude depending on the different sampling distribution for new task. This indicates adaptive sampling technique is important to make $\kappa$ solely depend on $d$. Finally, we provide empirical results of a noisy grid-world environment to corroborate our theoretical findings.
DOI: --
发表时间: 2019-06
期刊: --
影响因子: --
作者:
M. Khodak;Maria-Florina Balcan;Ameet Talwalkar
通讯作者: M. Khodak;Maria-Florina Balcan;Ameet Talwalkar