Provable General Function Class Representation Learning in Multitask Bandits and MDPs

Provable General Function Class Representation Learning in Multitask Bandits and MDPs
复制标题

DOI:
10.48550/arxiv.2205.15701
复制
发表时间:
2022-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Rui Lu;Andrew Zhao;S. Du;Gao Huang
Rui Lu;Andrew Zhao;S. Du;Gao Huang
中科院分区:
其他
文献类型:
--
作者:
Rui Lu;Andrew Zhao;S. Du;Gao Huang

文献摘要

相似文献

虽然多任务表征学习已经成为强化学习(RL)中提高样本效率的一种流行方法,但对它为什么以及如何工作的理论理解仍然有限。由于分析一般函数类的表示遇到了诸如泛化保证、抽象函数空间中置信界的确定等非平凡的技术障碍,以往的分析工作只能假设表示函数是智能体已知的或线性函数类已知的。然而,线性情况分析严重依赖于线性函数类的特殊性,而现实世界的实践通常采用一般的非线性表示函数,如神经网络。这大大降低了它的适用性。在这项工作中,我们将分析扩展到一般函数类表示。具体来说,我们认为代理播放$M$上下文土匪(或MDP)同时提取一个共享的表示函数$\phi$从一个特定的功能类$\Phi$使用我们提出的广义功能上界置信算法(GFUCB)。我们第一次从理论上验证了多任务表示学习在一般函数类中对强盗和线性MDP的好处。最后,我们进行实验,以证明我们的算法与神经网络表示的有效性。
While multitask representation learning has become a popular approach in reinforcement learning (RL) to boost the sample efficiency, the theoretical understanding of why and how it works is still limited. Most previous analytical works could only assume that the representation function is already known to the agent or from linear function class, since analyzing general function class representation encounters non-trivial technical obstacles such as generalization guarantee, formulation of confidence bound in abstract function space, etc. However, linear-case analysis heavily relies on the particularity of linear function class, while real-world practice usually adopts general non-linear representation functions like neural networks. This significantly reduces its applicability. In this work, we extend the analysis to general function class representations. Specifically, we consider an agent playing $M$ contextual bandits (or MDPs) concurrently and extracting a shared representation function $\phi$ from a specific function class $\Phi$ using our proposed Generalized Functional Upper Confidence Bound algorithm (GFUCB). We theoretically validate the benefit of multitask representation learning within general function class for bandits and linear MDP for the first time. Lastly, we conduct experiments to demonstrate the effectiveness of our algorithm with neural net representation.