Representation and generalization in autonomous reinforcement learning

Representation and generalization in autonomous reinforcement learning
复制标题

DOI:
10.14279/depositonce-5715
复制
发表时间:
2017
期刊:
--
影响因子:
--
通讯作者:
Wendelin Böhmer
Wendelin Böhmer
中科院分区:
其他
文献类型:
--
作者:
Wendelin Böhmer

文献摘要

相似文献

这篇博士论文研究了自主强化学习(RL)中泛化价值观的表征的作用。我采用函数分析的数学框架来检查连续(或混合)状态和动作空间中的归纳和演绎强化学习方法。我的分析揭示了表征度量对于归纳概括的重要性,以及结构假设将值演绎归纳到全新情况的必要性。该论文对两个相关的研究领域做出了贡献:大型状态动作空间中的表示学习和演绎价值估计。我在这里强调代理的自主权,要求对可能的任务和环境知之甚少甚至不了解(或限制)。下面我更详细地总结了我的贡献。我认为所有同构状态空间(以及完全可观察的观察空间)仅在其度量上有所不同(Böhmer 等人,2015),并且表明可以通过扩散度量来最佳地概括值。我证明,当通过最小二乘策略迭代(LSPI)等线性算法估计值时,慢特征分析(SFA)近似于同一环境中所有任务的最佳表示(Böhmer 等人,2013)。我通过推导一种新颖的正则化稀疏核 SFA 算法(RSK-SFA,Böhmer 等人,2012)来证明这一主张,并将学习到的表示与其他表示进行比较,例如在真实的 LSPI 机器人导航任务及其广泛的模拟中。我还将 SFA 的定义扩展到 γ-SFA,它仅代表预期任务的特定子集。自主归纳学习在许多现实任务中都受到样本不足的困扰,因为由许多变量组成的环境可能具有指数级的多种状态。这排除了归纳表示学习以及这些环境的归纳和演绎价值估计,所有这些都需要足够接近每个状态的训练样本。我提出了关于状态动力学的结构性假设来打破诅咒。我研究具有稀疏条件独立转换的状态空间,称为贝叶斯动态网络 (DBN)。与价值函数不同,稀疏 DBN 转换模型可以归纳学习,而不会遭受上述诅咒。为此,我定义了一类新的线性因子函数(LFF、Böhmer 和 Obermayer,2015),它可以对整个函数进行分析计算 DBN、边缘化和逐点乘法中的运算。我推导了保持 LFF 紧凑的压缩算法和三种归纳 LFF 算法密度估计、回归和值估计。由于归纳值估计受到样本不足的困扰,我推导了 LSPI 的演绎变体(FAPI,Böhmer 和 Obermayer,2013)。与 LSPI 一样,FAPI 需要预先定义的基函数,因此无法在大型状态动作空间中自主估计值。因此,我开发了第二种算法,直接在 LFF 的函数空间中演绎地估计 DBN(由 LFF 表示,例如,通过 LFF 回归学习)的值。由于大多数环境无法通过 DBN 完美建模,因此我讨论了一种将归纳和演绎价值估计相结合的重要性采样技术。此外,可以通过混合专家 DBN 进一步改进推导,其中一组条件确定每个状态的专家最能描述动态。条件可以根据基础关系规则构建为 LFF。这允许为每个环境配置生成一个转换模型。最终,本文得出的框架可以将归纳学习模型推广到具有类似对象的其他环境。
This PhD thesis investigates the role of representations that generalize values in autonomous reinforcement learning (RL). I employ the mathematical framework of function analysis to examine inductive and deductive RL approaches in continuous (or hybrid) state and action spaces. My analysis reveals the importance of the representationŠs metric for inductive generalization and the need for structural assumptions to generalize values deductively to entirely new situations. The thesis contributes to two related Ąelds of research: representation learning and deductive value estimation in large state-action spaces. I emphasize here the agentŠs autonomy by demanding little to no knowledge about (or restrictions to) possible tasks and environments. In the following I summarize my contributions in more detail. I argue that all isomorphic state spaces (and fully observable observation spaces) difer only in their metric (Böhmer et al., 2015), and show that values can be generalized optimally by a difusion metric. I prove that when the value is estimated by a linear algorithm like least squares policy iteration (LSPI), slow feature analysis (SFA) approximates an optimal representation for all tasks in the same environment (Böhmer et al., 2013). I demonstrate this claim by deriving a novel regularized sparse kernel SFA algorithm (RSK-SFA, Böhmer et al., 2012) and compare the learned representations with others, for example, in a real LSPI robot-navigation task and in extensive simulations thereof. I also extend the deĄnition of SFA to γ-SFA, which represents only a speciĄed subset of ŞanticipatedŤ tasks. Autonomous inductive learning sufers the curse of insufficient samples in many realistic tasks, as environments consisting of many variables can have exponentially many states. This precludes inductive representation learning and both inductive and deductive value estimation for these environments, all of which need training samples suiciently ŞcloseŤ to every state. I propose structural assumptions on the state dynamics to break the curse. I investigate state spaces with sparse conditional independent transitions, called Bayesian dynamic networks (DBN). In diference to value functions, sparse DBN transition models can be learned inductively without sufering the above curse. To this end, I deĄne the new class of linear factored functions (LFF, Böhmer and Obermayer, 2015), which can compute the operations in a DBN, marginalization and point-wise multiplication, for an entire function analytically. I derive compression algorithms to keep LFF compact and three the inductive LFF algorithms density estimation, regression and value estimation. As inductive value estimation sufers the curse of insuicient samples, I derive a deductive variant of LSPI (FAPI, Böhmer and Obermayer, 2013). Like LSPI, FAPI requires predeĄned basis functions and can thus not estimate values autonomously in large state-action spaces. I develop therefore a second algorithm to estimate values deductively for a DBN (represented by LFF, e.g., learned by LFF regression) directly in the function space of LFF. As most environments can not be perfectly modeled by DBN, I discuss an importance sampling technique to combine inductive and deductive value estimation. Deduction can be furthermore improved by mixture-of-expert DBN, where a set of conditions determines for each state which expert describes the dynamics best. The conditions can be constructed as LFF from grounded relational rules. This allows to generate a transition model for each conĄguration of the environment. Ultimately, the framework derived in this thesis could generalize inductively learned models to other environments with similar objects.