课题基金 / 基金详情

Adaptive Learning With Indirect Payoff Information

Adaptive Learning With Indirect Payoff Information
具有间接收益信息的自适应学习
批准号:
0111781
负责人:
Dana Heller
金额:
$2.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2001
资助国家:
美国
项目状态:
已结题
起止时间:
2001-07-01 至 2003-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
适应性学习指的是一个过程,在这个过程中,代理从他们对行动的直接经验中学习到他们行动的价值。有证据表明,适应性学习的变体很好地描述了受试者在实验中的行为。这个项目调查代理从自己的经验中适应性地学习做出的决定,也从他们没有采取的行动的间接信息中学习。特别是,我们调查了这种间接信息的扭曲对决策质量的影响,我们将其解释为对信息来源的态度。我们发现,行为模式取决于代理人相对于客观收益是夸大还是压低间接收益信息。如果代理夸大间接收益信息,有一段时间没有玩的操作往往看起来比实际情况更好,导致代理偶尔重新访问它们。结果,行为通过不止一个动作“循环”。另一方面,如果代理人倾向于减少间接信息,他会决定扮演一个单一的行动。由于未发挥作用的替代方案被认为比它们的客观情况更差,所以代理人认为是最好的限制行动在客观上可能是次优的。前一种模式类似于求变,而后一种模式类似于“锁定”行为。市场研究表明,这两种模式在消费者的选择中都可以观察到。学习规则的第二个关键特征是对过去观察结果的贴现程度。特别是,上期收益的权重是否变得任意变小。我们提供了关于贴现作用的初步结果,表明贴现的效果关键取决于是否利用间接收益信息。虽然折扣会导致在固定环境中的次优行为,但我们会在涉及有限的决策情况下寻找理由。
英文摘要
Adaptive learning refers to a process where agents learn about the value of their actions from their direct experience with the actions. There is evidence that variants of adaptive learning well describe subjects' behavior in experiments. This project investigates decisions made by agents learning adaptively from their own experience but also from indirect information about actions they did not take. In particular, we investigate the effect that distortions of such indirect information, which we interpret as representing attitudes towards the source of information, have on the quality of the decisions. We find that the pattern of behavior depends on whether the agent inflates or deflates indirect payoff information relative to the objective payoffs. If the agent inflates indirect payoff information, actions that have not been played for a while tend to look better than they truly are, leading the agent to revisit them every so often. As a result, behavior ``cycles'' through more than one action. On the other hand, if the agent tends to deflate indirect information, he will settle to playing a single action. Since the unplayed alternatives are perceived to be worse than they objectively are, the limit action, which the agent perceives to be the best, can be objectively suboptimal. The former pattern resembles variety seeking while the latter resembles ``lock-in'' behavior. Marketing research suggests that both patterns are observed in consumer choice. The second key feature of the learning rule is the degree to which past observations are discounted. In particular, whether the weight put on last period's payoff becomes arbitrarily small or not. We provide preliminary results on the role of discounting, suggesting the effect depends crucially on whether indirect payoff information is utilized or not. While discounting leads to suboptimal behavior in a stationary environment, we look for justifications for it in decision situations which involve limited
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: