Learning to Efficiently Plan in Flexible Distributed Organizations
Learning to Efficiently Plan in Flexible Distributed Organizations
批准号:
EP/R001227/2
负责人:
Frans Oliehoek
金额:
$5.2万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2018
资助国家:
英国
项目状态:
已结题
起止时间:
2018 至 --
中文摘要
机器人团队有望给工业和社会其他领域带来革命性的变化。然而,在这种所谓的多智能体系统(MASS)中,不确定条件下的决策在计算上是非常复杂的。分散部分可观测马尔可夫决策过程(DEC-POMDP)框架有助于对此类决策问题的原则性描述,但目前还没有可扩展的求解方法来保证任务性能。为了简化大规模的协调,代理组织向每个代理分配一个抽象的、更容易的问题。通常情况下,只有最严格的组织才能带来明显的计算效益,这些组织完全分离了代理。然而,这些都是以任务性能为代价的:完全分离意味着代理不再能够协作来分配工作负载。这个项目将专注于DEC-POMDP的灵活分布式组织(FDO),它将考虑的交互限制在空间附近的代理上,而不实施完全解耦。目前,对于FDO还没有可扩展的、能保证任务绩效的决策方法:该项目的主要目标是开发这种方法以及支持其形式化的理论。为了实现这一目标,它将研究使用深度学习技术来学习FDO中“影响”的表示,并使用这些表示来开发新的规划方法。如果成功,这将提供概念证明,即学习的影响力表征可以使大规模的原则性决策成为可能。这将是一个更大的研究计划的基础,该研究计划调查不同形式抽象的这种影响表示,并将引发应用研究,调查开发的算法在真实机器人团队中的部署。
英文摘要
Teams of robots are expected to revolutionise industry and other other parts of society. However, decision making in such so-called multiagent systems (MASs) under uncertainty is computationally very complex. The decentralized partially observable Markov decision process (Dec-POMDP) framework facilitates principled formulation of such decision making problems, but currently there are no scalable solution methods that provide guarantees on task performance. To simplify coordination in MASs, agent organisations assign an abstracted, easier problem to each agent. Typically only the most rigid organisations, which completely decouple the agents, have led to clear computational benefits. However, these come at the expense of task performance: full decoupling means that agents can no longer collaborate to divide the workload. This project will focus on flexible distributed organisations (FDOs) for Dec-POMDPs, which restrict considered interactions to spatially nearby agents without imposing full decoupling. Currently no scalable decision making methods with guarantees on task performance exist for FDOs: the main goal of the project is to develop such methods along with the theory that supports their formalisation. To accomplish this goal, it will investigate the use of deep learning techniques to learn representations of 'influence' in FDOs and use those representations to develop novel planning methods. If successful, this will provide the proof-of-concept that learned influence representations can enable principled decision making in large-scale MASs. This will be the basis for a larger research program investigating such influence representations for different forms of abstraction and will spark applied research that investigates deployment of the developed algorithms in real robotic teams.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
Analysing Factorizations of Action-Value Networks for Cooperative Multi-Agent Reinforcement Learning
分析协作多智能体强化学习的行动价值网络分解
DOI:
10.48550/arxiv.1902.07497
发表时间:
2019
期刊:
arXiv e-prints
影响因子:
--
作者:
[Castellini Jacopo]
通讯作者:
Castellini Jacopo
Beyond Local Nash Equilibria for Adversarial Networks
超越对抗性网络的局部纳什均衡
DOI:
10.48550/arxiv.1806.07268
发表时间:
2018
期刊:
arXiv e-prints
影响因子:
--
作者:
[Oliehoek Frans A.]
通讯作者:
Oliehoek Frans A.
DOI:
--
发表时间:
2018-11
期刊:
影响因子:
--
作者:
[Sammie Katt;F. Oliehoek;Chris Amato]
通讯作者:
Sammie Katt;F. Oliehoek;Chris Amato
DOI:
10.24963/ijcai.2020/12
发表时间:
2020-03
期刊:
影响因子:
--
作者:
[A. Czechowski;F. Oliehoek]
通讯作者:
A. Czechowski;F. Oliehoek
DOI:
10.24963/ijcai.2018/813
发表时间:
2018-07
期刊:
影响因子:
--
作者:
[F. Oliehoek]
通讯作者:
F. Oliehoek
Learning to Efficiently Plan in Flexible Distributed Organizations
-
批准号:EP/R001227/1
-
项目类别:Research Grant
-
资助金额:$12.88万
-
财政年份:2017
-
负责人:Frans Oliehoek
-
依托单位:
海外基金