Contingency, Contiguity, and Causality in Conditioning: Applying Information Theory and Weber's Law to the Assignment of Credit Problem

Contingency, Contiguity, and Causality in Conditioning: Applying Information Theory and Weber's Law to the Assignment of Credit Problem
复制标题

DOI:
10.1037/rev0000163
复制
发表时间:
2019-10-01
影响因子:
5.4
通讯作者:
Shahan, Timothy A.
Shahan, Timothy A.
中科院分区:
心理学1区
文献类型:
--
作者:
Gallistel, C. R.;Craig, Andrew R.;Shahan, Timothy A.

文献摘要

被引文献

相似文献

偶然性是关联学习理论和在强化学习中分配信用问题的关键概念。但是,测量和操纵它是有问题的。偶然性相互信息介绍的信息理论定义它很容易计算出强化事件之间关系的属性,预测它们的刺激及其产生它们的响应。必要时。所需的时间表示的动态范围除以韦伯的分数,给出了熵的心理现实插件估计。当鸽子按照可变的增强间隔时间表时,啄和加固之间没有可衡量的前瞻性偶然性。但是,在加固和紧接的佩克之间存在完美的回顾性意外。免费加强的回顾性偶然性降低了临界价值(.25),在下面的性能迅速下降。偶然性是时间尺度不变的,而对近端因果关系的看法取决于我假设存在很短。固定在心理上可以忽略的因果关系之间的关键间隔。增加了响应与增强之间的间隔,使其触发逆行的偶然性降低,导致绩效下降,使其恢复到其临界值或高于其临界值。因此,加强的回顾性效果没有关键的间隔。最后,我们对信息理论偶然性的广泛解释范围进行了简短的综述,当时是条件中的因果变量。我们建议,意外情况的计算可以取代强化学习模型中所有未来奖励的总和的计算。
Contingency is a critical concept for theories of associative learning and the assignment of credit problem in reinforcement learning. Measuring and manipulating it has, however, been problematic. The information-theoretic definition of contingency-normalized mutual information-makes it a readily computed property of the relation between reinforcing events, the stimuli that predict them and the responses that produce them. When necessary. the dynamic range of the required temporal representation divided by the Weber fraction gives a psychologically realistic plug-in estimates of the entropies. There is no measurable prospective contingency between a peck and reinforcement when pigeons peck on a variable interval schedule of reinforcement. There is, however, a perfect retrospective contingency between reinforcement and the immediately preceding peck. Degrading the retrospective contingency by gratis reinforcement reveals a critical value (.25), below which performance declines rapidly. Contingency is time scale invariant, whereas the perception of proximate causality depends-we assume-on there being a short. fixed psychologically negligible critical interval between cause and effect. Increasing the interval between a response and reinforcement that it triggers degrades the retrograde contingency, leading to a decline in performance that restores it to at or above its critical value. Thus, there is no critical interval in the retrospective effect of reinforcement. We conclude with a short review of the broad explanatory scope of information-theoretic contingencies when regarded as causal variables in conditioning. We suggest that the computation of contingencies may supplant the computation of the sum of all future rewards in models of reinforcement learning.