Contrasting temporal difference and opportunity cost reinforcement learning in an empirical money-emergence paradigm.

Contrasting temporal difference and opportunity cost reinforcement learning in an empirical money-emergence paradigm.
复制标题

DOI:
10.1073/pnas.1813197115
复制
发表时间:
2018-12-04
影响因子:
11.1
通讯作者:
Palminteri S
Palminteri S
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Lefebvre G;Nioche A;Bourgeois-Gironde S;Palminteri S

文献摘要

参考文献

被引文献

相似文献

在本研究中,我们将实验经济学中未经典使用的强化学习模型应用于货币出现的多步骤交换任务,该任务源自货币出现的经典搜索理论范式。这种方法使我们能够强调机会成本反事实反馈处理在投机性使用货币的学习过程中的重要性,以及强化学习模型对多步骤经济任务的预测能力。这些结果向理解在多步骤经济决策中起作用的学习过程和金钱使用的认知微观基础迈出了一步。在现代经济中,货币是一种基本的、无处不在的制度。然而,它的出现问题仍然是经济学家的核心问题。货币搜索理论方法研究商品货币作为一种解决方案出现的条件,以克服分散经济中个体间交换固有的摩擦。尽管在这些条件中,主体的理性是经典的必要条件,也是任何理论货币均衡的先决条件,但当执行搜索理论范式的任务时,当这些策略是投机性的,即涉及使用昂贵的交换媒介来增加后续和成功交易的概率时,人类主体往往无法采用最优策略。在目前的工作中,我们假设实现这种投机行为依赖于强化学习,而不是经典经济理论所假设的终身效用计算。为了验证这一假设,我们将Kiyotaki和Wright在多步骤交换任务中的货币出现范式进行了操作,并将人类受试者执行该任务的行为数据与两个强化学习模型进行了拟合。他们中的每一个都实现了一个关于当前决策中未来或反事实奖励权重的独特认知假设。我们发现,这两个模型都优于理论预测的受试者关于投机策略的实施行为,后者依赖于学习过程中考虑机会成本的程度。因此,对货币的市场性优势的推测似乎依赖于代理人在交换情况下对反事实事件的心理模拟。
In the present study, we applied reinforcement learning models that are not classically used in experimental economics to a multistep exchange task of the emergence of money derived from a classic search-theoretic paradigm for the emergence of money. This method allowed us to highlight the importance of counterfactual feedback processing of opportunity costs in the learning process of speculative use of money and the predictive power of reinforcement learning models for multistep economic tasks. Those results constitute a step toward understanding the learning processes at work in multistep economic decision-making and the cognitive microfoundations of the use of money. Money is a fundamental and ubiquitous institution in modern economies. However, the question of its emergence remains a central one for economists. The monetary search-theoretic approach studies the conditions under which commodity money emerges as a solution to override frictions inherent to interindividual exchanges in a decentralized economy. Although among these conditions, agents’ rationality is classically essential and a prerequisite to any theoretical monetary equilibrium, human subjects often fail to adopt optimal strategies in tasks implementing a search-theoretic paradigm when these strategies are speculative, i.e., involve the use of a costly medium of exchange to increase the probability of subsequent and successful trades. In the present work, we hypothesize that implementing such speculative behaviors relies on reinforcement learning instead of lifetime utility calculations, as supposed by classical economic theory. To test this hypothesis, we operationalized the Kiyotaki and Wright paradigm of money emergence in a multistep exchange task and fitted behavioral data regarding human subjects performing this task with two reinforcement learning models. Each of them implements a distinct cognitive hypothesis regarding the weight of future or counterfactual rewards in current decisions. We found that both models outperformed theoretical predictions about subjects’ behaviors regarding the implementation of speculative strategies and that the latter relies on the degree of the opportunity costs consideration in the learning process. Speculating about the marketability advantage of money thus seems to depend on mental simulations of counterfactual events that agents are performing in exchange situations.
强化学习是情绪低落的有条件合作行为:实验结果。
DOI: 10.1038/srep39275
发表时间: 2017-01-10
期刊: Scientific reports
影响因子: 4.6
作者:
Horita Y;Takezawa M;Inukai K;Kita T;Masuda N
通讯作者: Masuda N
DOI: 10.1093/rfs/14.1.1
发表时间: 2001-03-01
影响因子: 8.2
作者:
Gervais, S;Odean, T
通讯作者: Odean, T
DOI: 10.1093/scan/nsl025
发表时间: 2006-12-01
影响因子: 4.2
作者:
Delgado, M. R.;Labouliere, C. D.;Phelps, E. A.
通讯作者: Phelps, E. A.
DOI: 10.1073/pnas.0608842104
发表时间: 2007-05-29
影响因子: 11.1
作者:
Lohrenz, Terry;McCabe, Kevin;Montague, P. Read
通讯作者: Montague, P. Read
DOI: 10.1126/science.1094550
发表时间: 2004-05-21
期刊: SCIENCE
影响因子: 56.9
作者:
Camille, N;Coricelli, G;Sirigu, A
通讯作者: Sirigu, A