课题基金 / 基金详情

MIMIc: Multimodal Imitation Learning in MultI-Agent Environments

MIMIc: Multimodal Imitation Learning in MultI-Agent Environments
MIMIc:多代理环境中的多模式模仿学习
批准号:
EP/T000783/1
负责人:
Varuna De Silva
金额:
$32.99万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2019
资助国家:
英国
项目状态:
已结题
起止时间:
2019 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在英国,17岁之前不允许开车。这是因为,驾驶是一项复杂而安全的关键活动,需要许多先进的认知技能,如识别可能的威胁,预测其他道路使用者的行为以及对新兴情况的敏捷反应。想想一个足球运动员在球场上做决定。一个好的球员可以通过预测其他球员会做什么来感知机会,并选择一个能增加得分几率的动作。人类需要很长时间来发展这些高级认知技能,成为如此复杂的现实世界任务的专家。人工智能在过去十年中取得了重大进展,癌症检测,计算机击败“围棋”大师和智能机器人技术的突破证明了这一点。然而,如果人工智能要实现其科幻小说中的承诺,帮助人类甚至取代人类智能,它至少应该具备人类所拥有的认知技能。该项目旨在开发突破性的算法,使自主系统具备在真实的世界环境中蓬勃发展所需的人类认知技能。我们专注于需要自主代理(例如机器人或无人驾驶汽车)与环境中的多个智能代理进行交互以完成任务的应用(称为多代理环境:MAE)。这样的应用程序需要一个代理预测其他代理的行为,并选择最合适的行动方案。使代理具有这种自主决策能力被称为策略学习。与单代理域中的策略学习(教机器人走路或教计算机玩视频游戏)相比,MAES中的策略学习最近的进展相当温和。这是由于多个原因:1)由于代理动作,环境是动态的2)多代理策略学习受到称为维度灾难(CoD)的理论限制3)捕获代理目标的效用函数难以定义4)严重缺乏足够的多代理数据集,允许有意义的研究。本项目拟通过克服上述局限性,对农业和经济实体的政策学习进行研究。我们独特的方法来政策学习MAES是由人类如何在类似的环境中茁壮成长的动机。首先,我们通过多种感官(即视觉,听觉,触觉)感知世界,从而对世界有丰富的感知。其次,当在MAE中行动时,人类不会注意所有的刺激,而是只注意关键的刺激,例如,当足球运动员进攻球时,球员只注意能够实现目标的队友和关键防守队员。最后,我们采用的学习范式被称为模仿学习,是一种通过观察专家来学习的新兴方法,这是我们用来学习新技能的一种富有成效的方法。因此,我们建议通过模仿学习,利用多模态数据融合和选择性注意力建模来学习MAES中的现实政策。多模态数据融合允许捕获真实的世界的高维上下文,并且选择性注意模型允许减轻CoD的问题。我们已经提供了一个独特的多模式多智能体数据集和访问国家的最先进的设施来捕捉数据,由一个精英足球俱乐部促进这个雄心勃勃的研究项目。项目输出将被主观验证作为一种工具来回答“如果”的问题,与足球比赛相关的帮助教练组可视化投机性的比赛策略,并作为计算基准来量化足球运动员的认知技能。计划中的影响活动将确保该项目将在人工智能开发方面留下遗产,通过在无人驾驶汽车、视频游戏和辅助机器人等多个高增长领域做出重大贡献,使英国PLC受益。
英文摘要
In UK, we are not allowed to drive a vehicle until we are 17. It is because, driving is a complex and safety critical activity that requires many advanced cognitive skills like recognition of possible threats, anticipation of behavior of other road users and agile reaction to emerging situations. Think about a football player making decisions on field. A good player can sense the opportunities, through anticipating what other players will do, and select an action that will increase the odds of scoring. It takes a long time for humans to develop these advanced cognitive skills, to become an expert at such complex real-world tasks. Artificial Intelligence has made significant progress during the last decade, demonstrated by breakthroughs in cancer detection, computers beating 'Go' masters and intelligent robotics. However, if AI is to live up to its science fictional promises to assist humanity or even supersede human intelligence, it should at least be equipped with cognitive skills such as those possessed by humans. This project aims to develop ground breaking algorithms that equip autonomous systems with human like cognitive skills required to thrive in real world environments.We are focused on applications that require autonomous agents (e.g. Robot or Driverless car) to interact with multiple intelligent agents in the environment to accomplish a task (known as Multi-Agent Environments: MAEs). Such applications require an agent to anticipate the behaviour of other agents and to select the most appropriate course of actions. Equipping agents with such autonomous decision-making capability is known as policy learning. Compared to policy learning in single agent domains (teaching a robot to walk or a computer to play a video game), the recent progress of policy learning in MAEs has been quite modest. This is due to multiple reasons: 1)Due to agent actions the environment is dynamic 2)multi-agent policy learning suffers from a theoretical limitation known as curse of dimensionality (CoD) 3)Utility functions that capture agent objectives are difficult to define 4)there is a significant lack of adequate multi-agent datasets that allow meaningful research. This project proposes to undertake research in to policy learning in MAEs, by addressing the above limitations. Our unique approach to policy learning in MAEs is motivated by how humans thrive in similar settings. Firstly, we perceive the world through multiple senses, (i.e. vision, audition, touch) enabling a rich perception of the world. Secondly, when acting in a MAE, humans do not pay attention to all the stimuli but only to key stimuli e.g. when a football player is attacking the ball, the player pays attention only to the teammates capable of effecting a goal and the key defenders. Finally, the learning paradigm we employ known as imitation learning is an emerging methodology to learn by observing experts, which is a productive approach that we use to learn new skills. Accordingly, we propose to learn realistic policies in MAEs through imitation learning by leveraging multimodal data fusion and selective-attention modelling. Multimodal data fusion allows to capture high dimensional context of the real world and selective attention model allows for allaying the issue of CoD. We have been provided a unique multimodal multi-agent dataset and access to state-of-the-art facilities to capture data, by an elite football club facilitating this ambitious research project.The project outputs will be subjectively validated as a tool to answer "what-if" questions related to game play in football assisting coaching staff to visualize speculative game strategies, and as a computational benchmark to quantify cognitive skills of football players. The planned impact activities will ensure the project will leave a legacy in AI development benefiting UK PLC through significant contribution in multiple high growth areas, such as driverless vehicles, video gaming, and assistive robots.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
Learning data-driven decision-making policies in multi-agent environments for autonomous systems
在自治系统的多代理环境中学习数据驱动的决策策略
DOI: 10.1016/j.cogsys.2020.09.006
发表时间: 2021
期刊: Cognitive Systems Research
影响因子: 3.9
作者: [Hook J]
通讯作者: Hook J
DOI: 10.3390/rs11232723
发表时间: 2019-11
期刊: Remote. Sens.
影响因子: --
作者: [Katie Inder;V. D. Silva;Xiyu Shi]
通讯作者: Katie Inder;V. D. Silva;Xiyu Shi
A machine learning framework for quantifying in-game space-control efficiency in football
用于量化足球比赛中空间控制效率的机器学习框架
DOI: 10.1016/j.knosys.2023.111123
发表时间: 2024
期刊: Knowledge-Based Systems
影响因子: 8.8
作者: [Gu C]
通讯作者: Gu C
Intelligent Systems and Pattern Recognition - Third International Conference, ISPR 2023, Hammamet, Tunisia, May 11-13, 2023, Revised Selected Papers, Part II
智能系统和模式识别 - 第三届国际会议,ISPR 2023,突尼斯哈马马特,2023 年 5 月 11-13 日,修订后的精选论文,第二部分
DOI: 10.1007/978-3-031-46338-9_12
发表时间: 2024
期刊:
影响因子: --
作者: [Artaud C]
通讯作者: Artaud C
海外基金