Shifting Inductive Bias with Success-Story Algorithm, Adaptive Levin Search, and Incremental Self-Improvement

Shifting Inductive Bias with Success-Story Algorithm, Adaptive Levin Search, and Incremental Self-Improvement
复制标题

DOI:
10.1023/a:1007383707642
复制
发表时间:
1997-07
期刊:
影响因子:
7.5
通讯作者:
J. Schmidhuber;Jieyu Zhao;M. Wiering
J. Schmidhuber;Jieyu Zhao;M. Wiering
中科院分区:
计算机科学3区
文献类型:
--
作者:
J. Schmidhuber;Jieyu Zhao;M. Wiering

文献摘要

被引文献

相似文献

我们研究通过适当的归纳偏见转变(学习者政策的改变)来加速学习者平均奖励摄入的任务序列。为了评估偏见转变的长期影响,为以后的偏见转变奠定基础,我们使用“成功故事算法”(SSA)。SSA有时会被调用,这可能取决于策略本身。它使用回溯来消除那些尚未被经验观察到的偏差变化,以触发长期奖励加速(直到当前SSA调用)。在SSA中幸存下来的偏见转变代表了一生的成功历史。在下一次SSA调用之前,它们被认为是有用的,并为额外的偏倚偏移奠定了基础。SSA允许插入各种各样的学习算法。我们插入(1)一个新的,自适应扩展的莱文搜索和(2)嵌入学习者的政策修改策略的政策本身(增量自我改进)的方法。我们的感应传输案例研究涉及复杂的,部分可观察的环境,传统的强化学习失败了。
We study task sequences that allow for speeding up the learner's average reward intake through appropriate shifts of inductive bias (changes of the learner's policy). To evaluate long-term effects of bias shifts setting the stage for later bias shifts we use the “success-story algorithm” (SSA). SSA is occasionally called at times that may depend on the policy itself. It uses backtracking to undo those bias shifts that have not been empirically observed to trigger long-term reward accelerations (measured up until the current SSA call). Bias shifts that survive SSA represent a lifelong success history. Until the next SSA call, they are considered useful and build the basis for additional bias shifts. SSA allows for plugging in a wide variety of learning algorithms. We plug in (1) a novel, adaptive extension of Levin search and (2) a method for embedding the learner's policy modification strategy within the policy itself (incremental self-improvement). Our inductive transfer case studies involve complex, partially observable environments where traditional reinforcement learning fails.