Adaptive Learning With Indirect Payoff Information
Adaptive Learning With Indirect Payoff Information
批准号:
0111781
负责人:
Dana Heller
金额:
$2.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2001
资助国家:
美国
项目状态:
已结题
起止时间:
2001-07-01 至 2003-05-31
中文摘要
自适应学习是指一个过程,在这个过程中,代理人从他们的直接经验中了解他们的行动的价值。有证据表明,适应性学习的变体很好地描述了实验中受试者的行为。该项目调查的决策,从自己的经验,但也从间接的信息,他们没有采取行动的代理学习自适应。特别是,我们调查的影响,这种间接的信息,我们解释为代表的态度,对信息源的扭曲,对决策的质量。我们发现,行为模式取决于代理人是否膨胀或收缩相对于客观回报的间接回报信息。如果代理人夸大了间接收益信息,那么一段时间没有采取的行动往往看起来比实际情况更好,导致代理人经常重新审视它们。因此,行为通过不止一个动作“循环”。另一方面,如果施动者倾向于减少间接信息,他将满足于采取单一行动。由于未采取的方案被认为比客观上更糟糕,因此代理人认为是最好的限制行动可能是客观上次优的。前一种模式类似于多样化寻求,而后者类似于“锁定”行为。 市场研究表明,这两种模式在消费者的选择中都可以观察到。学习规则的第二个关键特征是过去的观察被打折的程度。特别是,上一期收益的权重是否任意变小。我们提供了初步的结果贴现的作用,这表明效果取决于是否利用间接收益信息或没有。虽然折扣导致次优行为在一个固定的环境中,我们寻找理由,它在决策的情况下,涉及有限的
英文摘要
Adaptive learning refers to a process where agents learn about the value of their actions from their direct experience with the actions. There is evidence that variants of adaptive learning well describe subjects' behavior in experiments. This project investigates decisions made by agents learning adaptively from their own experience but also from indirect information about actions they did not take. In particular, we investigate the effect that distortions of such indirect information, which we interpret as representing attitudes towards the source of information, have on the quality of the decisions. We find that the pattern of behavior depends on whether the agent inflates or deflates indirect payoff information relative to the objective payoffs. If the agent inflates indirect payoff information, actions that have not been played for a while tend to look better than they truly are, leading the agent to revisit them every so often. As a result, behavior ``cycles'' through more than one action. On the other hand, if the agent tends to deflate indirect information, he will settle to playing a single action. Since the unplayed alternatives are perceived to be worse than they objectively are, the limit action, which the agent perceives to be the best, can be objectively suboptimal. The former pattern resembles variety seeking while the latter resembles ``lock-in'' behavior. Marketing research suggests that both patterns are observed in consumer choice. The second key feature of the learning rule is the degree to which past observations are discounted. In particular, whether the weight put on last period's payoff becomes arbitrarily small or not. We provide preliminary results on the role of discounting, suggesting the effect depends crucially on whether indirect payoff information is utilized or not. While discounting leads to suboptimal behavior in a stationary environment, we look for justifications for it in decision situations which involve limited
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Understanding structural evolution of galaxies with machine learning
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:Nicola Rosario Napolitano
-
依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:吉建娇
-
依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
-
批准号:62003314
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:沈剑
-
依托单位:
集成上下文张量分解的e-learning资源推荐方法研究
-
批准号:61902016
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2019
-
负责人:万珊珊
-
依托单位:
具有时序迁移能力的Spiking-Transfer learning (脉冲-迁移学习)方法研究
-
批准号:61806040
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2018
-
负责人:解修蕊
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于有向超图的大型个性化e-learning学习过程模型的自动生成与优化
-
批准号:61572533
-
项目类别:面上项目
-
资助金额:66.0万元
-
批准年份:2015
-
负责人:孙雪冬
-
依托单位:
E-Learning中学习者情感补偿方法的研究
-
批准号:61402392
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2014
-
负责人:秦继伟
-
依托单位: