Value-free reinforcement learning: policy optimization as a minimal model of operant behavior

Value-free reinforcement learning: policy optimization as a minimal model of operant behavior
复制标题

DOI:
10.1016/j.cobeha.2021.04.020
复制
发表时间:
2021-05-28
影响因子:
5
通讯作者:
Langdon, Angela J.
Langdon, Angela J.
中科院分区:
心理学2区
文献类型:
--
作者:
Bennett, Daniel;Niv, Yael;Langdon, Angela J.

文献摘要

被引文献

相似文献

强化学习是一个强大的框架,用于建模学习和决策的认知和神经基础。认知神经科学和神经经济学的当代研究通常使用基于价值的认知学习模型,该模型假设决策者通过比较不同行为的学习价值来进行选择。然而,另一种可能性是由一个更简单的模型家族提出的,称为策略梯度强化学习。策略梯度模型通过直接优化行为策略来学习,而不需要中间的价值学习步骤。在这里,我们回顾了最近的行为和神经的研究结果,更简约的解释政策梯度模型比基于价值的模型。我们的结论是,尽管无处不在的“价值”的决策学习模型,政策梯度模型提供了一个轻量级的和令人信服的替代模型的操作性行为。
Reinforcement learning is a powerful framework for modelling the cognitive and neural substrates of learning and decision making. Contemporary research in cognitive neuroscience and neuroeconomics typically uses value-based reinforcement-learning models, which assume that decision-makers choose by comparing learned values for different actions. However, another possibility is suggested by a simpler family of models, called policy-gradient reinforcement learning. Policy-gradient models learn by optimizing a behavioral policy directly, without the intermediate step of value-learning. Here we review recent behavioral and neural findings that are more parsimoniously explained by policy-gradient models than by value-based models. We conclude that, despite the ubiquity of 'value' in reinforcement-learning models of decision making, policy-gradient models provide a lightweight and compelling alternative model of operant behavior.