Near-Optimal Scalable Algorithms for Multi-Agent Reinforcement Learning
Near-Optimal Scalable Algorithms for Multi-Agent Reinforcement Learning
批准号:
2444539
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --
中文摘要
序贯决策是现代统计理论和应用的一个重要背景,其中智能体顺序地与环境交互-观察其状态,采取行动并获得奖励-以最大化累积奖励为目标。这类方法已成功应用于游戏、金融、机器人、自动驾驶和计算机视觉。这些应用程序中的许多涉及多个代理的参与,这就带来了状态和动作空间的可扩展性问题。事实上,虽然许多已经建立的算法实现了良好的缩放保证单代理设置,他们不执行,如果他们立即应用到多个代理在同一环境中进行交互的情况下。这是因为最简单的方法是将状态或动作视为所有代理的联合状态或动作,从而导致空间相对于代理的数量呈指数级增长。我们的方法包括在划分成邻域的代理集,以减少指数缩放的setting.The项目的目的是获得一个接近最佳的算法的多代理设置的计算复杂性,指数只依赖于一个邻域的基数。该项目属于EPSRC“统计学与应用概率”和“人工智能技术”研究领域。福尔斯。
英文摘要
Sequential decision making is an important setting of modern statistical theories and applications, where an agent sequentially interacts with an environment - observing its state, taking actions and receiving rewards - with the objective of maximizing the cumulative reward. This class of methods has been successfully applied to games, finance, robotics, autonomous driving and computer vision. Many of these applications involve the participation of multiple agents, which brings a scalability issue with respect to the state and action spaces. Indeed, while many of the already established algorithms achieve good scaling guarantees for the single agent setting, they do not perform well if they are immediately applied to the case where multiple agents interact in the same environment. This is because the simplest approach would be to consider a state or an action as the joint states or actions of all the agents, causing the spaces to grow exponentially with respect to the number of agents. Our approach consists in dividing the set of agents into neighbourhoods in order to reduce the exponential scaling of the setting.The aim of the project is to obtain a near-optimal algorithm for the multi-agent setting with a computational complexity that depends exponentially only on the cardinality of a neighbourhood. This project falls within the EPSRC "Statistics an applied probability" and "Artificial intelligence technologies" research areas.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金