课题基金 / 基金详情

Adaptive Learning With Indirect Payoff Information

Adaptive Learning With Indirect Payoff Information
具有间接收益信息的自适应学习
批准号:
0111781
负责人:
Dana Heller
金额:
$2.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2001
资助国家:
美国
项目状态:
已结题
起止时间:
2001-07-01 至 2003-05-31

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
适应性学习指的是一个过程,在这个过程中,智能体从他们对行动的直接经验中了解他们的行动的价值。有证据表明,适应性学习的变体很好地描述了实验对象的行为。这个项目调查了代理人从他们自己的经验中适应性学习做出的决定,也从他们没有采取的行动的间接信息中学习。特别是,我们研究了这种间接信息的扭曲对决策质量的影响,我们将其解释为对信息来源的态度。我们发现,行为模式取决于代理是否夸大或缩小间接支付信息相对于客观支付。如果代理夸大了间接收益信息,那么一段时间未执行的操作就会看起来比实际情况更好,从而导致代理经常重新访问它们。因此,行为通过不止一个动作“循环”。另一方面,如果代理倾向于减少间接信息,他将采取单一行动。由于未发挥作用的备选方案被认为比它们客观上的情况更糟,因此代理认为最好的极限行动客观上可能是次优的。前一种模式类似于寻求多样性,而后者类似于“锁定”行为。市场研究表明,这两种模式在消费者的选择中都可以观察到。学习规则的第二个关键特征是对过去观察的折扣程度。特别是,上一时期收益的权重是否变得任意小。我们提供了关于贴现作用的初步结果,表明其影响关键取决于是否利用了间接支付信息。虽然折扣在固定环境中导致次优行为,但我们在涉及有限的决策情况下寻找它的理由
英文摘要
Adaptive learning refers to a process where agents learn about the value of their actions from their direct experience with the actions. There is evidence that variants of adaptive learning well describe subjects' behavior in experiments. This project investigates decisions made by agents learning adaptively from their own experience but also from indirect information about actions they did not take. In particular, we investigate the effect that distortions of such indirect information, which we interpret as representing attitudes towards the source of information, have on the quality of the decisions. We find that the pattern of behavior depends on whether the agent inflates or deflates indirect payoff information relative to the objective payoffs. If the agent inflates indirect payoff information, actions that have not been played for a while tend to look better than they truly are, leading the agent to revisit them every so often. As a result, behavior ``cycles'' through more than one action. On the other hand, if the agent tends to deflate indirect information, he will settle to playing a single action. Since the unplayed alternatives are perceived to be worse than they objectively are, the limit action, which the agent perceives to be the best, can be objectively suboptimal. The former pattern resembles variety seeking while the latter resembles ``lock-in'' behavior. Marketing research suggests that both patterns are observed in consumer choice. The second key feature of the learning rule is the degree to which past observations are discounted. In particular, whether the weight put on last period's payoff becomes arbitrarily small or not. We provide preliminary results on the role of discounting, suggesting the effect depends crucially on whether indirect payoff information is utilized or not. While discounting leads to suboptimal behavior in a stationary environment, we look for justifications for it in decision situations which involve limited
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Understanding structural evolution of galaxies with machine learning
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2022
  • 负责人:
    Nicola Rosario Napolitano
  • 依托单位:
煤矿安全人机混合群智感知任务的约束动态多目标Q-learning进化分配
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    吉建娇
  • 依托单位:
基于领弹失效考量的智能弹药编队短时在线Q-learning协同控制机理
  • 批准号:
    62003314
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    沈剑
  • 依托单位: