The Successor Representation and Temporal Context

The Successor Representation and Temporal Context
复制标题

DOI:
10.1162/neco_a_00282
复制
发表时间:
2012-06-01
期刊:
影响因子:
2.9
通讯作者:
Sederberg, Per B.
Sederberg, Per B.
中科院分区:
计算机科学4区
文献类型:
--
作者:
Gershman, Samuel J.;Moore, Christopher D.;Sederberg, Per B.

文献摘要

被引文献

相似文献

后继表示由Dayan(1993)引入到强化学习中,作为促进具有相似后继的状态之间的泛化的一种手段。尽管强化学习总体上已被广泛用作心理和神经过程的模型,但后继表示的心理有效性尚未得到探索。一个有趣的可能性是,后继表示不仅可以用于强化学习,也可以用于情景学习。我们的主要贡献是表明,时间上下文模型(TCM;霍华德& Kahana,2002)的一个变体,一个有影响力的情景记忆模型,可以被理解为直接估计的继任者表示使用时间差异学习算法(萨顿& Barto,1998)。这种见解导致了中医的推广和新的实验预测。除了为中医药提供新的规范外,这种等效性还表明不同学习系统之间存在以前未探索过的接触点。
The successor representation was introduced into reinforcement learning by Dayan (1993) as a means of facilitating generalization between states with similar successors. Although reinforcement learning in general has been used extensively as a model of psychological and neural processes, the psychological validity of the successor representation has yet to be explored. An interesting possibility is that the successor representation can be used not only for reinforcement learning but for episodic learning as well. Our main contribution is to show that a variant of the temporal context model (TCM; Howard & Kahana, 2002), an influential model of episodic memory, can be understood as directly estimating the successor representation using the temporal difference learning algorithm (Sutton & Barto, 1998). This insight leads to a generalization of TCM and new experimental predictions. In addition to casting a new normative light on TCM, this equivalence suggests a previously unexplored point of contact between different learning systems.