课题基金 / 基金详情

Epistemic Uncertainty Estimation in Multi-Agent Reinforcement Learning

Epistemic Uncertainty Estimation in Multi-Agent Reinforcement Learning
多智能体强化学习中的认知不确定性估计
批准号:
2747642
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
该项目属于EPSRC人工智能(AI)和机器人研究领域。强化学习(RL)是一种通过奖励系统训练AI代理的技术。智能体被允许在环境中进行交互和执行动作,当智能体成功地完成预定任务时就会获得奖励。RL有着广泛的应用,包括机器人、推荐系统和医疗保健。多智能体强化学习(MARL)关注的是多个智能体如何在共同的环境中相互作用。每个智能体都受到自己的奖励和兴趣的激励,但智能体可以协作以实现共同的目标或相互竞争,从而产生复杂的群体动态。对MARL的研究越来越重要,随着人工智能代理在我们日常生活的许多方面(例如自动驾驶汽车)变得广泛。目标和目标在现实世界中,代理对周围世界没有完美的知识是常见的,而建模的不确定性是避免灾难性和危险故障的根本。这意味着,代理人应该知道它不知道的东西。这个项目的目标是提供对Marl环境中认知不确定性的准确估计。认知不确定性指的是由于缺乏知识而导致的不确定性。在Marl方案中,这包括有关环境或其他代理人的动机和行为的信息。这种不确定性可以通过采取行动探索环境或与其他智能体交互来减少。获得正确和校准的不确定性估计可能会导致智能体之间更安全的交互和协作。至关重要的是,这包括人工智能代理与人类之间的互动。这一互动的相关应用将是自动驾驶汽车与人类司机互动,或使用不同软件的其他自动驾驶汽车。对所有其他代理的行为进行不确定性建模,无论他们运行什么软件或他们的意图是什么,都是有效和安全合作的基础。相比之下,贝叶斯模型提供了一个理论上扎根的框架来推理模型的不确定性,但由于其极高的计算成本,通常不可能在除最简单的环境之外的所有环境中使用。最近,已经提出了多种技术来规避这一挑战和近似贝叶斯推理,例如神经网络中的辍学(Gal,2016)和深度集成(Lakshminarayanan,2017)。“作为贝叶斯近似的丢弃:代表深度学习中的模型不确定性。”机器学习国际会议。Lakshminarayanan,Balaji,Alexander Pritzel和Charles Blundell。使用深度集成进行简单且可扩展的预测不确定性评估。神经信息处理系统进展30(2017)。
英文摘要
This project falls in the EPSRC artificial intelligence (AI) and robotics research area.Reinforcement Learning (RL) is a technique to train an AI agent through a system of rewards. The agent is allowed to interact and execute actions in an environment, and rewards are awarded when the agent successfully completes the intended task.RL has a multitude of applications, some examples include robotics, recommendations systems and healthcare.Multi-Agent Reinforcement Learning (MARL) focuses on how multiple agents interact with each other in a common environment.Each agent is motivated by their own rewards and interest, but agents can collaborate to achieve common goals or compete with each other, resulting in complex group dynamics.The study of MARL is increasingly more relevant, as AI agents become widespread in many aspects of our daily lives (e.g. self-driving cars).Aims and ObjectivesIn real-world scenarios, it is common for agents to not have perfect knowledge of the world around them, and modelling uncertainty is fundamental to avoid catastrophic and dangerous failures. This means, the agent should know what it does not know.The goal of this project is providing an accurate estimate of epistemic uncertainty in the MARL setting.Epistemic uncertainty refers to uncertainty caused by a lack of knowledge. In the MARL scenario, this includes information about the environment or other agents' motivations and behaviour. This type of uncertainty can be reduced by taking actions to explore the environment or interact with other agents.Obtaining a correct and calibrated uncertainty estimate could lead to safer interactions and collaboration between agents. Crucially, this includes interactions between AI agents and humans.A relevant application of this would be self-driving cars interacting with human drivers or other self-driving cars using different software. Modelling uncertainty over all other agents' behaviours, regardless of what software they run or what their intentions are, is fundamental for effective and safe collaboration.Novelty of the research methodologyCurrent RL techniques are incredibly successful, but fail to model uncertainty. In contrast, Bayesian models offer a theoretically grounded framework to reason about model uncertainty, but are often impossible to use in all but the simplest environments, due to their extremely high computational costs. Recently, multiple techniques have been proposed to circumvent this challenge and approximate Bayesian inference, such as dropout in Neural Networks (Gal, 2016) and Deep Ensembles (Lakshminarayanan, 2017).Gal, Yarin, and Zoubin Ghahramani. "Dropout as a bayesian approximation: Representing model uncertainty in deep learning." international conference on machine learning. PMLR, 2016.Lakshminarayanan, Balaji, Alexander Pritzel, and Charles Blundell. "Simple and scalable predictive uncertainty estimation using deep ensembles." Advances in neural information processing systems 30 (2017).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金