Belief and Truth in Hypothesised Behaviours

Belief and Truth in Hypothesised Behaviours
复制标题

DOI:
10.1016/j.artint.2016.02.004
复制
发表时间:
2015-07
期刊:
Artif. Intell.
影响因子:
--
通讯作者:
Stefano V. Albrecht;J. Crandall;S. Ramamoorthy
Stefano V. Albrecht;J. Crandall;S. Ramamoorthy
中科院分区:
其他
文献类型:
--
作者:
Stefano V. Albrecht;J. Crandall;S. Ramamoorthy

文献摘要

被引文献

相似文献

博弈论中关于贝叶斯或“理性”学习的主题有着悠久的历史,其中每个参与者都对其他参与者的一组替代行为或类型保持信念。这个想法在人工智能(AI)社区中引起了越来越多的兴趣,在那里它被用作一种方法来控制由具有未知行为的多个代理组成的系统中的单个代理。这个想法是假设一组类型,每个类型指定其他代理的可能行为,并根据我们认为最有可能的类型计划我们自己的行动,给定代理的观察行为。博弈论文献主要在均衡实现的背景下研究这一思想。相比之下,许多人工智能应用程序都专注于任务完成和回报最大化。考虑到这一点,我们确定并解决了一系列与假设类型中的信念和真理有关的问题。我们制定了三种基本的方法,将证据纳入后验信念,并显示当由此产生的信念是正确的,当他们可能不正确。此外,我们证明了先验信念可以对我们在长期内最大化收益的能力产生重大影响,并且它们可以自动计算,具有一致的性能效果。此外,我们分析的条件下,我们能够最佳地完成我们的任务,尽管不准确的假设类型。最后,我们展示了如何通过自动统计分析在交互过程中确定假设类型的正确性。
There is a long history in game theory on the topic of Bayesian or “rational” learning, in which each player maintains beliefs over a set of alternative behaviours, or types, for the other players. This idea has gained increasing interest in the artificial intelligence (AI) community, where it is used as a method to control a single agent in a system composed of multiple agents with unknown behaviours. The idea is to hypothesise a set of types, each specifying a possible behaviour for the other agents, and to plan our own actions with respect to those types which we believe are most likely, given the observed actions of the agents. The game theory literature studies this idea primarily in the context of equilibrium attainment. In contrast, many AI applications have a focus on task completion and payoff maximisation. With this perspective in mind, we identify and address a spectrum of questions pertaining to belief and truth in hypothesised types. We formulate three basic ways to incorporate evidence into posterior beliefs and show when the resulting beliefs are correct, and when they may fail to be correct. Moreover, we demonstrate that prior beliefs can have a significant impact on our ability to maximise payoffs in the long-term, and that they can be computed automatically with consistent performance effects. Furthermore, we analyse the conditions under which we are able complete our task optimally, despite inaccuracies in the hypothesised types. Finally, we show how the correctness of hypothesised types can be ascertained during the interaction via an automated statistical analysis.