Conditions and methods for Decentralised Reinforcement Learning
Conditions and methods for Decentralised Reinforcement Learning
批准号:
2619847
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
在过去的十年中,强化学习(RL)的发展使其成为研究的热点。硬件性能的改进以及RL与神经网络的使用相结合,使得算法的开发在许多控制问题上实现了最先进的性能,包括他们击败人类冠军的计算机游戏。然而,该领域仍然存在的一些悬而未决的问题是如何在更复杂的环境中学习,如何从有限的样本中更有效地学习,以及如何为更一般的任务学习。在非常复杂的环境中学习的一种方法是将控制任务分散到多个代理,而不是单一的集中式代理。这可以极大地降低每个代理人学习的复杂性,但可能会花费一组代理人可以制定的更有限的政策(行动计划),以及影响培训过程稳定性的其他技术问题。在许多情况下,去中心化非常自然地发生,例如在自动驾驶汽车中,其中每个代理可以是一辆车,或者在计算机集群中的资源分配任务中,其中每个代理可以控制分配给每台计算机的任务。这些分散的代理可以彼此具有不同级别的通信和同步,这影响了代理将采取的可能策略集的大小。我的研究旨在处理隐式通信的代理,这意味着它们不直接相互共享状态(状态信息),但它们观察环境的共同特征,允许收集有关其他代理状态的信息。我打算回答的第一个问题是,在什么情况下,这样一个分散的RL系统可以达到与集中式单一代理相同的性能水平?这涉及到对州、奖励和政策设定数学条件。这是通过将分散的解建模为分散的部分可观测的马尔可夫决策过程(DEC-POMDP)来实现的,该过程允许考虑代理的分散性和来自每个代理的环境的部分可观测性。然后,我想在更一般的场景中进行调查,应用去中心化的效果是什么?在某些条件下,我可以推导出去中心化造成的性能损失的理论界限吗?在某些特殊情况下,去中心化特别有用吗?随后,我想使用这些条件来开发一种算法,它可以很容易地区分哪些任务是可分散的,哪些任务不是。根据数学条件的不同,直接使用导出的公式可以很容易地做到这一点,但它也可能涉及大量的计算。在这种情况下,创建近似值将是有用的,这将允许容易地测试分散的问题任务的解如何影响理论性能界限。虽然现有的许多关于分散的RL算法试图在每种情况下获得最大性能的研究,但集中的和分散的RL解决方案之间的关系还没有被深入地探索。我的博士研究旨在提供关于这种关系的理论基础,并旨在以算法的形式提供新的工具,使分散解决方案的设计者能够知道分散解决方案的特定设计可以实现的最大理论性能。这项研究有可能应用于许多系统的控制,这些系统具有需要协作行为以实现最佳性能的部件。这类系统的例子可以在自动驾驶汽车、机器人、通信网络等领域找到。我的研究与ESPRC领域的“人工智能技术”和“ICT网络和分布式系统”保持一致。
英文摘要
Advances in Reinforcement learning (RL) in the last decade have made it a hot topic for research. Improvements in hardware performance and the combination of RL with the use of neural networks have allowed for the development of algorithms that achieve state-of-the-art performance in many control problems, including computer games in which they beat human champions. Some open questions that remain in the field however are how to learn in more complex environments, how to learn more efficiently from limited samples and how to learn for more general tasks. One approach used to learn in very complex environments is to decentralise the control task to multiple agents, rather than a single centralised one. This can greatly reduce the complexity of learning by each agent with the possible expense of more limited policies (action plans) that can be enacted by the group of agents and other technical issues affecting the stability of the training process. The decentralisation occurs quite naturally in many scenarios, such as in self-driving cars, in which each agent can be one car or in a resource assignment task in a cluster of computers, in which each agent could control the tasks assigned to each computer. These decentralised agents can have various levels of communication and synchronisation with each other, which affects the size of the set of possible policies to be taken by the agents. My research aims to deal with agents that communicate implicitly, meaning that they do not directly share their status (state information) with each other, however they observe common features of the environment that allow to collect information about the status of the other agents. The first question that I aim to answer is what are the scenarios in which such a decentralised RL system can achieve the same level of performance as a centralised single agent? This involves setting mathematical conditions on the states, the rewards, and the policy. This is done by modelling the decentralised solution as a decentralised partially observable Markov decision process (dec-POMDP), which allows to consider the decentralisation of the agents and the partial observability of the environment from each agent. Then, I want to investigate in more general scenarios, what is the effect of applying decentralisation? Can I derive theoretical bounds on the performance loss due to decentralisation under certain conditions? Are there special conditions under which decentralisation is especially useful? Subsequently, I want to use these conditions to develop an algorithm that can easily distinguish between tasks that are decentralizable and those that are not. Depending on what the mathematical conditions are, this may be easily done directly using the derived formula, but it could also involve massive computation. In this case, it would be useful to create approximations that would allow to easily test how decentralising the solution of the problem task affect the theoretical performance bounds.While there has been much existing research about decentralised RL algorithms trying to achieve the maximum performance in every kind of scenario, the relationship between centralised and decentralised RL solutions has not been explored in depth. My PhD research aims to provide a theoretical foundation about this relationship and aims to provide novel tools in the form of algorithms that would allow the designer of a decentralised solution to know the maximum theoretical performance that a certain design of a decentralised solution can achieve. This research has the potential to be applied to the control of many systems having components that require cooperating behaviour to achieve the optimal performance. Examples of such systems can be found in self-driving cars, robotics, communication networks, etc. My research is aligned with the ESPRC field "Artificial Intelligence technologies" and "ICT networks and distributed systems".
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
复杂图像处理中的自由非连续问题及其水平集方法研究
-
批准号:60872130
-
项目类别:面上项目
-
资助金额:28.0万元
-
批准年份:2008
-
负责人:刘国才
-
依托单位:
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: